d91beef0fd314250d8d9b94de86dfea019a8bd96
46
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a95041f1e6 |
feat(studio): speech-to-speech voice changer (#1765)
Add a bounded, local-first speech-to-speech Convert workflow with shared ASR/TTS admission, duration matching, stale-request cancellation, profile conditioning, watermarking, persistence, localization, and regression coverage.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com> |
||
|
|
4053397921 |
feat(dub): karaoke word-highlight caption burn-in (#1764)
Adds opt-in word-timed ASS karaoke captions while preserving the existing line-caption default.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com> |
||
|
|
497d57ee62 |
Show complete engine disk costs before install (#1728)
Closes #1718. Adds structured pre-install and measured post-install disk costs, complete sidecar preflight accounting, strict authorization for recursive scans, and localized catalogue details with confidence and deduplication context. |
||
|
|
b92d35ac5d |
feat: add local speech platform (#1671)
Fixes #1646 |
||
|
|
5a615d2c66 |
feat(workers): package headless GPU nodes (#1638) (#1648)
Closes #1638.\n\nPackages headless GPU workers with durable enrollment, bounded artifact handling, cross-platform lifecycle cleanup, and regression coverage. Incorporates CodeRabbit, Greptile, CodeQL, and platform-CI findings before merge. |
||
|
|
48c9a3b1f8 |
feat(settings): compute-device override (auto / CUDA / ROCm / XPU / MPS / CPU) (#1557)
* feat(settings): compute-device override — auto | CUDA | ROCm | XPU | MPS | CPU Auto-detect stays the default; the override kills the 'auto-detect picked wrong' issue class. Applied at the single choke point (_probe()'s family selection) so routing, get_best_device(), and every badge inherit it. Resolution: OMNIVOICE_DEVICE env > Settings pick (prefs.json) > auto (#981 pattern). An override can steer, never invent hardware: a family the host lacks is noted and ignored; cpu is always honorable. Applies at next backend start (host caps are immutable per process — same restart contract as the rest of the Performance tab, RestartBadge shown). GET/PUT /api/settings/compute-device (admin-gated) reports resolved vs applied so the panel shows restart-required truthfully and disables itself under an env pin instead of pretending. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): entry for the compute-device override (#1557) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(device-override): harvest — the override reaches CT2 ASR, full i18n, honest edge states - _ctranslate2_cuda_ok() and the ASR sidecar now gate on the probe's family, so a cpu pin (or ROCm host) can never hand CTranslate2 a CUDA device — the override reaches every CT2 loader through one shared gate - override_ignored exposed by the API and shown by the panel (env pin naming a device this machine lacks: auto is in effect, restart won't change it) - all 8 panel strings + 5 device-family labels translated into all 21 locales; failed saves keep their error visible through the re-sync - test isolation: cleanup drops OMNIVOICE_DEVICE before re-probing so no overridden caps leak into later tests; panel tests wait for loaded state - xpu/intel search keywords; oxfmt formatting Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(device-override): round 2 — fail-safe probe fallbacks, complete i18n, combined pin state - a broken capability probe now means CPU everywhere (CT2 gate + ASR sidecar) — never a torch-derived guess that would bypass a cpu pin or re-open #1529 on ROCm; regression test added - env-pinned AND not-detected shows both facts in one subtitle - device_load_failed/perf_save_failed translated into all 21 locales; CJK/th/vi/ar strings no longer say literal 'Auto' - test_ctranslate2_never_gets_cuda_on_a_rocm_build pins the probe family (it was order-dependent on the lru_cache before) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): pin the probe family in the faster-whisper OOM-fallback test Same class as the rocm-build test: it mocked torch but not the probe the new override gate consults first, so on a cpu-family CI host the CUDA fallback chain under test was unreachable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bb813ff676 |
feat(startup): bind the socket in ~1s and narrate startup step by step (#1550)
* feat(startup): bind the socket in ~1s and narrate startup step by step The structural fix for the "can't reach the local backend" class (~1 in 5 of every issue ever filed): uvicorn served nothing until torch import (10-20s cold), the 30-router fan-out, an import-time DB migration, the cuDNN preload, and alembic all finished — every slow or fragile step rendered as an unexplained dead backend. main.py now keeps module scope fast and defers the heavy work: - _phase_a_build (executor thread): prefs/env restore + #963 migration, yt-dlp overlay, cuDNN preload, torchaudio, model_manager, router imports — order preserved, literal imports so PyInstaller still traces. - _phase_a_finalize (event loop, no awaits → atomic wrt requests): include_router, mounts, MCP, SPA, openapi bust. - _phase_b: the old lifespan startup body; handles on app.state so shutdown survives a startup that never finished. - Eager mode (pytest / OMNIVOICE_EAGER_INIT=1) runs everything at import — byte-equivalent behavior for the ~100 lifespan-less TestClient sites and for embedders (dump_api_routes, probe boot runner opt in). While starting: /health answers 503 with the current step, new /startup/progress serves the full ledger (always 200), and StartupGateMiddleware 503s everything else with the [starting] marker (same skip-the-Report-button convention as [shutting_down]). A deferred failure keeps import-crash semantics: traceback to stderr → shell crash forensics, run sentinel stays uncleared, exit 1 names the failed step. Shell: startup_progress() probe (marker-header-gated so a foreign responder can't narrate the splash) feeds per-step log lines into the launch poll and the supervisor's reconnect wait. --health-check absorbs the deferred init (60→180s); --diagnose runs Phase A up front so it still sees restored prefs. Docker HEALTHCHECK semantics unchanged (curl -f fails on 503 exactly as it did on connection-refused). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(startup): join the Phase A thread on shutdown; async fail-path sleep Bot-review harvest on #1550: cancelling the deferred-startup task cannot stop the executor thread inside Phase A's blocking imports — shutdown now waits (bounded, only when a build started and hasn't finished) on a thread-completion event so interpreter teardown can't race a mid-import (#1000 class). The failure path's last-poll beat is now awaited, not time.sleep — a blocking sleep froze the very loop that beat exists to let serve. Also: dump_api_routes forces eager (assignment, not setdefault), and the integration test's child gets DEVNULL instead of an undrained pipe that could wedge a cold boot. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(startup): close the Phase A submission race; CodeQL nits Review finds on #1550: shutdown could sample _phase_a_started unset while the executor callable was queued-but-not-running, skipping the thread join. started is now set BEFORE submission, the submission is shielded so a cancel can't strand a queued callable that would never set _phase_a_finished, and the wrapper sets finished on every exit including the already-built early return. Contract pinned by test_phase_a_thread_join_contract. Plus explanatory comments on the new bare excepts and a consistent return in the gate's websocket branch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0ee62b2261 |
feat(gallery): save gallery voices as profiles, with validated audio references (#1542)
* feat(gallery): save gallery voices as profiles, with validated audio references Work-in-progress lifted from the concurrent gallery session at the owner's request (its uncommitted working tree, preserved verbatim from base 92b1ee5d; safety snapshot remains at rescue/gallery-wip): - gallery voices can be saved as local profiles: audio is copied into the profile store with content-addressed filenames, existing profiles are detected and refreshed only when the source clip changed - backend/core/audio_validation.py: symlink-rejecting, root-contained resolution for persisted profile WAV references, with tests - archetype/community routers and the Voice Gallery UI updated for the save-as-profile handoff (spec: docs/specs/longform/26-gallery-use-handoff.md) - locale updates for the new gallery strings across all 21 files Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: drop a stray local screenshot script that rode in with the tree copy * fix(community): explain the tolerated Content-Length parse failure; drop an unused import CodeQL on #1542: the empty except now says why it is safe (the streamed byte counter enforces the same cap regardless), and the test file loses an unused Path import. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(gallery): review findings — copy outside the write lock, no stale completions CodeRabbit on #1542, all findings addressed: - the profile-audio copy stages to a .part temp BEFORE BEGIN IMMEDIATE and publishes via atomic os.replace inside it — other backend writers no longer block for the duration of an audio copy; a mid-copy failure leaves no temp droppings and no profile row (both pinned by tests) - VoiceGallery async ops carry per-operation generation tokens: a preview or save-as-profile that resolves after unmount (or after a newer operation) can no longer play audio, redirect into a workspace, or touch state — three fail-before regression tests - VoiceGalleryActions imports the page at test runtime; the e2e locator uses a stable data-testid instead of a translated string; symlink tests skip cleanly where the OS can't create symlinks; the changelog line carries its PR ref Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: static ffmpeg fallback when the chocolatey feed is down Third feed outage to break a PR run (2026-07-20, 2026-07-28, today — three attempts, three 'installed 0/1'). Chocolatey is a distribution channel, not the dependency: after the retry loop exhausts, fetch the static gyan.dev build from its GitHub release mirror and put it on PATH — same binary, no feed in the path. URL verified live (HTTP 200). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
19ae20111a |
fix(security): replace persistent admin keys with scoped sessions (#1528)
* fix(security): replace persistent admin keys with sessions Exchange the remote administrator key once for bounded, revocable credentials. Canonicalize backend principals, enforce cookie CSRF and exact origins, and use path-bound one-use WebSocket tickets. Migrate the bundled UI away from durable master-key storage and credential-bearing URLs. Add unit, integration, static-hygiene, and production-browser regressions plus synchronized operator documentation. * docs: link session hardening to PR 1528 * fix(security): key session indexes with process pepper Use HMAC-SHA-256 instead of an unkeyed digest for in-memory session and WebSocket-ticket indexes. This preserves constant-size lookup identifiers, makes copied records unusable without the process pepper, and resolves CodeQL's weak sensitive-data hash finding. * fix(auth): align empty bearer migration precedence Centralize the Authorization-channel presence decision with canonical principal parsing. Bearer followed only by spaces now remains an empty channel during legacy cookie migration, while unsupported or invalid explicit credentials stay authoritative and fail closed. * fix(security): harden admin session review boundaries * fix(security): derive key generations with HKDF * fix(auth): anchor the admin-session store so module reloads cannot fork it test_master_exchange_does_not_bypass_pin_on_normal_routes failed in full-suite runs: test_mcp_bindings' client fixture purges the services.* tree from sys.modules and reloads main, so api.routers.auth re-imported a fresh services.admin_sessions (new AdminSessionStore) while core.auth kept its import-time reference to the old one — the exchange issued the cookie into one store and the middleware resolved it against another, turning the expected "PIN required" into "API key required". Root cause is the class of bug, not the one test: a process-global auth store defined as a bare module-level singleton forks under importlib.reload or purge-and-reimport. Fix at the source: admin_session_store now resolves through a synthetic sys.modules anchor (_omnivoice_admin_session_store_anchor) that reloads never re-execute and package-prefix purges never match, so every copy of the module shares the one per-process store. No consumer or behavior changes. Regression test reproduces both fork vectors (in-place reload and sys.modules purge + fresh import) and asserts previously issued sessions still resolve and the store identity is preserved; it fails before this fix and passes after. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(auth): honor X-Forwarded-Proto for CSRF origin and Secure cookies behind TLS proxies Behind Tailscale Serve (docs/remote-gpu.md) or any TLS-terminating proxy, the browser talks https while the backend hop stays http, so exact-origin CSRF compared an https Origin against an http expectation and rejected every legitimate request, and the session cookie shipped without Secure. uvicorn's ProxyHeadersMiddleware only rewrites the scope for loopback peers, which misses Docker and any non-loopback proxy topology. New core.csrf.effective_scheme derives the client-facing scheme: resolved scope first (uvicorn's trusted-proxy rewrite wins), then an upgrade-only read of X-Forwarded-Proto's first value — https/wss promotes http to https, everything else is ignored, and a genuine TLS hop can never be downgraded. Used by both the destination-origin comparison and auth._secure_cookie so the WS-ticket/logout CSRF paths and the cookie Secure flag agree. Spoofing gains nothing: the host:port half of the origin tuple is untouched, browsers cannot attach the header cross-site without a preflight this API never grants, and forging it on plain http only adds Secure (the browser then drops the cookie — self-harm only). Regression tests: proxied https origin accepted (origin check, Secure flag, logout), comma-separated chains, scope-fallback path, spoofed header still rejects cross-origin, cannot downgrade real https, junk values ignored. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(auth): consume the stored admin key only after a successful exchange A remote-backend user upgrading with their backend unreachable lost the only stored copy of OMNIVOICE_API_KEY: every migration path deleted the durable ov_api_key BEFORE the session exchange settled, stranding them until they recovered the key from the server box. Close the whole class: - client.ts bootstrap: read the legacy key, exchange first, and remove the durable copy only after the exchange succeeds; on failure the key stays so the next launch retries the migration (auth gate still rises). - authSession.ts exchangeApiKey: move removeLegacyMaster from before the fetch to the cookie/bearer success paths — the key never coexists with a live session, but a rejected or hung exchange no longer consumes it. - remoteBackendProbe.ts configuredRemoteBackend: stop wiping the key on every app mount. - RemoteBackendPanel: a connection test or an aborted save no longer wipes the pending key; only disabling the remote backend discards it. - prefKeys.js: ov_api_key moves from PREF_KEYS to PRESERVED_KEYS — factory reset preserves the pending connection credential exactly like ov_backend_url; the successful migration is what deletes it. Tighten the credential-hygiene static guard to match: it accepted sessionStorage.setItem('ov_api_key', …) — the exact class it exists to close. The guard now flags .setItem(<master key>) on any storage receiver, quote style, or injected-store alias, with a self-test pinning what it catches and what stays legal. Fail-before/pass-after regression tests: backend unreachable retains the key and the next bootstrap retries it; a successful exchange removes it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * perf(auth): make session validation occupancy-independent * test(auth): catch optional master-key storage calls * feat(docs): add PR control document for bultodepapas in VoiceStudio * docs: keep the PR tracking board in the fork; credit the changelog line The pr-control document is excellent process discipline, but it is the contributor's own operational board (their inventory, their update commands) — it lives naturally in their fork, and docs/agents/ here is context every repo agent loads. Removed with appreciation; the changelog line gains its contributor credit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a310141114 |
feat(workers): join from the app, share by QR, and a status-bar Compute control (#1516)
* feat(workers): join from the app, share by QR, and a status-bar Compute control Remote workers shipped with a hole in the middle: the control plane could mint join codes, and on the other machine there was nothing to paste them into. Becoming a worker meant launching with OMNIVOICE_WORKER_MODE and OMNIVOICE_WORKER_TOKEN in the environment and relaunching — on the machine that is usually the least convenient one to configure by hand. Backend - GET /workers/agent, POST /workers/agent/join, POST /workers/agent/enabled. Join redeems a code and starts the agent live; no restart. - Worker mode now persists in settings as well as the environment (env still wins, and the panel is told so it can disable a switch it cannot honour), and it is written only after a join that actually worked — a failed enrolment must not have the app retrying on every launch. - The endpoint carried by the redeemed code is remembered. Without that a machine that joined from the UI came back up enrolled but with nowhere to dial, and the only fix was OMNIVOICE_WORKER_ENDPOINT. UI - "Lend this machine's GPU": paste the code, Join. Once joined it offers a switch rather than another code, because the pinned certificate survives. - <OneTimeSecret/> renders join codes and connection strings as a QR next to the text, with a live expiry countdown, and is used by both halves. QR generation is best-effort: a string past the format's capacity still shows the code and Copy, because losing the QR is a degraded share and losing the only copy of a one-time secret is data loss. - Status-bar Compute control: pick local or a machine, flip the feature, mint a join code — without opening Settings. Absent entirely until the user has opted in or enrolled something. - Remote workers now reads as a device list: status dot, address, latency, live task meter, resident models, last seen; housekeeping actions revealed on hover; a three-step empty state. - Approve is on the row. A worker could connect, sit there labelled "Not approved" and never be usable, with no way out of it in the UI. Fixes found on the way - Status dots and menu surfaces in the GPU picker were painted from fixed Tailwind palette classes (bg-emerald-400, text-amber-400, hover:bg-white/5), so on Midnight or Catppuccin they showed Gruvbox colours next to the theme's own. Both controls now paint from themed --color-* tokens, shared in computeTarget.jsx along with the JSON wrapper all three copies duplicated. - Button funnels every child into one <span>, so an icon passed as a child renders glued to its label — the flex gap only applies to the `leading` slot. Six buttons across these panels were affected. - InboundNodePanel passed `variant="warning"` to Badge, which takes `tone`; the "on your network" warning rendered as an ordinary neutral pill. Docs updated in the same change (docs/remote-workers.md): the join flow, the QR, the status-bar control, and the new environment variable. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(changelog): stamp the remote-workers entries with their PR ref Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(workers): a join is not done until the control plane accepts it Review findings on #1516: - Greptile P1: `start()` only SCHEDULES the dial-out loop, so a control plane that rejected this worker — expired token, wrong address, a server that never answers — looked identical to a successful join. The route persisted worker mode, reported success, and the machine retried forever on every launch. The agent now signals first registration, and join waits for it before persisting anything. - CodeRabbit: a failed REJOIN left the machine unable to reconnect to the control plane it was already serving, because pinning the new certificate overwrites the old one on disk. Snapshot the pinned certificate, endpoint and setting up front, and restore them (and the running agent) when the join fails. - CodeRabbit: join and the enable toggle awaited stop()/start() with no exclusion, so two concurrent requests could interleave their pairs and have `start()` return early — reporting success for a control plane it never dialled. Both now hold one lifecycle lock. - CodeRabbit: with OMNIVOICE_WORKER_MODE set, the toggle still started or stopped the agent and wrote a setting the rest of the app ignores, contradicting the env_pinned status it reports. It now answers 409 and says which variable is in charge. - CodeRabbit: the QR code kept encoding the previous secret until the new one finished encoding, so the code on screen could disagree with the text beside it. CI: regenerated tests/fixtures/api_routes.txt for the three /workers/agent routes. Tests: a join the control plane never accepts is a 409 that persists nothing and leaves no agent dialling; a failed rejoin restores the previous certificate, endpoint and setting; an env-pinned machine refuses the toggle. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(workers): the environment pin governs joining too, not just the toggle CodeRabbit, #1516: - join_control_plane skipped the env_pinned guard set_agent_enabled enforces, and joining is precisely what ENABLES worker mode: under OMNIVOICE_WORKER_MODE it wrote a setting nothing consults, and with the variable pinned off it handed back a machine that reported a successful join and lent nothing. One shared guard now covers both routes. - Two of the three rollback assertions could not fail before the fix (nothing wrote those settings on the failure path). The test now pins the behaviour only the rollback produces: the previous enrollment is dialling again, rather than left stopped until someone notices. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2b6f49c596 |
feat(workers): make inbound mode reachable — settings, endpoints, docs
Wires the two transport halves into something a user can actually turn on. Two independent switches, deliberately not one. "Accept connections" makes this machine a node others dial; "saved connections" are the nodes this panel dials out to. A workstation with a GPU that also drives jobs on a second box does both, so neither implies the other. Binding stays on 127.0.0.1 until someone explicitly widens it, and widening is its own field rather than a flag riding along with the enable toggle. With no encryption that boundary is the difference between a credential on one machine and a credential on a network, so it is never crossed as a side effect. The API reports `exposed` so the UI can say which side of it the user is on. Saved nodes are redialled only after the control plane is up, since the connector hands frames to its servicer. Failing to listen records the reason rather than leaving the feature looking enabled while it quietly accepts nothing. Docs say plainly that this mode is unencrypted, that the connection string is a password crossing the network in the clear, and that dial-out remains the better choice when one machine is enough. The Security section no longer implies its TLS guarantees cover both modes. |
||
|
|
b7caa494eb |
feat(workers): remote downloads, audiobook chapters, and one port that stays honest
Five workstreams that finish the remote-GPU line, plus the test hole that let a broken signature reach a commit. **Downloads go through the normal path** (Phase 5). Rather than a second remote-only route, the existing Models install flow became target-aware, so a model landing on a worker uses the same code, the same progress events and the same UI as a local one. Progress rows key on (target, repo_id) — the aggregator keyed on bare repo_id, so the same model downloading here and on a worker at once collapsed into one row that told the user nothing true about either. **Audiobooks render chapter by chapter on the worker** (Phase 8), with per-chapter local fallback and ONE aggregated notice. The failure that shape exists to prevent: a remote GPU that sleeps at chapter 40 of 200 must not turn a working book into 160 rows of PROGRESS_LEASE_EXPIRED. Dictation is deliberately NOT ported — it runs ASR per utterance inside a live WebSocket loop, and paying queue admission plus a round trip there would spend the one thing that route is for. **Dubbing stays local, and says so** (Phase 7). The coarse worker operation is not finished, so the picker still reports dubbing as local rather than showing a green remote chip over work this machine is doing. What could not wait is the in-loop OOM retry: it sniffed the error string and flushed the *local* CUDA cache, which under remote execution is the wrong machine's GPU entirely. That is fixed now, before the path that would have exercised it exists. **Two instances can no longer share the control plane.** A second VoiceStudio silently bound the same worker port and coexisted, so remote workers landed on whichever process won the race — a session that registers with one instance and appears dead to the other. This produced hours of misdiagnosis during hardware testing and would hit any user with the app open twice. The second instance now keeps running locally and explains the conflict instead of quietly competing. **And the hole that allowed all this to be missable.** gpu_gateway called Scheduler.submit(pinned_worker_id=...) one commit before that parameter existed. Every remote generation raised TypeError; 5236 tests passed anyway, because nothing exercised the gateway against the real scheduler. tests/test_gpu_gateway_scheduler_contract.py now runs that path for real and binds every gateway→dependency call signature. Verified by renaming the parameter away and watching both tests fail with the original error. Gallery previews also fall back to a local render when a downloaded clip cannot be decoded, rather than yielding silence. Backend 5274 passed, frontend 1812 passed. Not yet verified on hardware: Phases 4, 5, 6, 7, 8. Only the TTS path and its artifact transport have been proven on a real GPU. |
||
|
|
bda169c900 |
feat(workers): pin work to the chosen GPU, and say when its model is missing
Three phases that only make sense together: a job that names a worker, a worker that reports honestly what it can actually run, and the small defects that made both lie. **Pinning** (Phase 1). `pinned_worker_id` is now honoured in both places that choose a worker — `eligible_workers` and `select_worker` build independent lists, so applying it to one silently leaked work onto whichever machine was least busy. The pin persists across a restart via an additive column, deliberately not alembic (justified in the code, per the precedent already in db.py): quitting mid-render used to drop it without a word. `max_attempts=1` was rejected as the mechanism — it makes the FIRST failure terminal, including the penalty-free ones a stale advisory view produces routinely. Cancel now actually reaches the worker. `WorkerServicer.cancel` had zero callers, so cancelling released the slot while the GPU thread kept running, and a late result could resurrect the task as COMPLETED — `commit_result` assigned that state directly, bypassing the transition table where CANCELLED is terminal by construction. **Honest capabilities** (Phase 4). A worker now probes whether weights are actually present, and a job stops BEFORE dispatch with a typed 409 naming the model and the machine, instead of failing mid-task. The probe fails OPEN: `is_cached`/`cache_is_complete` cannot see a user-managed clone outside the HF layout, so only a positive "absent" refuses. Refusing an engine that works today would break the compatibility promise. `pool.supports` deliberately still ignores `downloaded` — had it not, the scheduler would drop the worker and answer with a terminal NO_CAPABLE_WORKER, which tells the user to check their install when the truth is one download away. The frontend no longer offers "Report this bug" for that state; it offers the download. Catalog tags resolve against the TARGET's OS/arch/backend, not this machine's. From a Mac control plane, a CUDA worker's model list was showing the mlx-community repos it cannot run and hiding the ones it needs. **And the quiet ones** (Phase 0 leftovers): a model's human label rides its own proto field so renaming it cannot orphan breaker history; an empty model_id no longer forks the capacity slot key into two slots for one model; the idle sweep cannot evict an engine out from under a live LOCAL render. Verified on real hardware, which is the only verification that has ever caught anything here: 2025 characters, default settings, routed to an RTX 4090 over the wire — 100% GPU utilisation on the remote box, 119.6 s of 24 kHz audio returned in 16.6 s, 5.7 MB delivered out of band through the artifact path rather than the control stream. Backend 5259 passed, frontend 1808 passed. |
||
|
|
c643706d07 |
feat(workers): make a remote GPU actually run a task, end to end
Selecting a remote worker repainted a badge and nothing else. The cause was not subtle: `scheduler.submit` had no production caller, and `routing.decide()` was read only by the status endpoint that paints the header. Remote execution was a complete, tested pipeline with no producer at its head. This adds the producer and fixes the defects that made the pipeline unable to carry a real job: - Nothing routed to the scheduler. Adds `POST /workers/tasks` (loopback-gated, **development-only** until the gateway lands) and `Scheduler.wait`, backed by per-task futures rather than the unregisterable `on_change` listener list. - Every task over two minutes died. No worker ever sent `TaskProgress`, so the 120s progress lease expired mid-render — including during the cold model load, which happens after `TaskStarted`. Workers now report progress and emit a keepalive, bounded by the phase's absolute budget so it renews the lease without deleting the only enforced bound in the system. - The executor rebuilt its engine per task (`return cls()`), so every job paid a cold load. Engines now share one instance cache with the router, resolved by the assignment's engine — never `get_active_tts_backend()`, which returns the worker machine's own Settings preference and would silently run the wrong engine. - One lease expiry took a worker offline permanently: parked slots were never reclaimed. Parks now expire on a TTL, and are deliberately NOT reconciled against the worker's own load report — at a ceiling of one the only task such a worker can report is the wedged one, so "busy" would drop the park and the next idle heartbeat would hand out a slot with a live GPU thread (#730/#1190). - A worker that dropped and reconnected mid-render had every liveness frame discarded: task frames were fenced on the live session epoch, which bumps on every reconnect, while the worker echoes the ref stamped at dispatch. The control plane then expired a task whose GPU was still rendering, and swallowed the failure report when it went wrong. Fenced per attempt instead. - A result from one worker could commit another's task, after which the owner's real delivery arrived as a duplicate and its audio was discarded. "Unknown attempt" and "another worker's attempt" are no longer the same answer. - An oversized result was a poison pill, re-sent identically on every reconnect and permanently disconnecting the worker. It is now a terminal `RESULT_TOO_LARGE`, which is also classified — it was falling through to TRANSIENT and retrying a re-render that could never fit. - `_store_inline` joined the artifact directory with worker-supplied ids, and `os.path.join` discards its prefix on an absolute component. Paths are now minted control-plane-side and resolved through `core.path_security`. - Remote synthesis bypassed `mark_synthetic`, and the guard that exists to catch exactly that walked only `backend/api` and `backend/services` — so it stayed green while a fourth unmarked producer shipped. Marking moved to the worker's tensor stage; the guard now walks `backend/worker` too. Also adds pre-rendered voice previews (`services/gallery.py`), so browsing the gallery no longer needs a GPU or a downloaded model. The manifest is verified against the updater's release key already baked into the binary; a fresh install hears voices without downloading 2.4GB first, and everything falls back to local rendering when the gallery is unreachable. Verified on hardware, not just in CI: 1728 characters submitted to an RTX 4090 returned 105.94s of 24kHz audio in 23.9s, committed and served from the artifact store. Not yet done, and deliberately not claimed: the keepalive fix cannot be exercised end-to-end on fast hardware, because any job long enough to reach the 120s lease produces audio past the 8MiB inline cap. Chunked `UploadResult` has to land first. Pinning to the worker the user chose is also still absent, so "Remote" reaches a remote GPU but not necessarily the one on the badge. |
||
|
|
9eb1ec7591 |
feat(workers): choose where jobs run, and show whether that machine is well
Adds a GPU target picker to the header: Local, or one of the machines you enrolled. Exactly one is active at a time; other connected workers are standby and receive nothing. The selection is the user's, not the scheduler's. The engine underneath can rank many workers and the hosted platform will need that, but a desktop app is better served by a choice you can predict and explain: "your worker is offline, this ran locally" is a sentence, "least-busy ranking preferred the laptop" is not. Picking an offline machine is allowed on purpose — you choose your desktop, then go and switch it on. `routing.decide()` is the single answer to "where does the next job run", shared by the badge and (soon) the generation path, so the badge cannot claim something the router will not do. It shows the RESOLVED answer rather than the stored choice: pick your desktop, let it sleep, and the chip reads Local with the reason, while the menu still shows your desktop selected. Connection latency is now real. `latency_ms` existed but nothing measured it — the protocol had Ping with no reply — so it was always zero. Adds Pong (additive, field 12) and times the round trip on the control plane's MONOTONIC clock, so an NTP step or a sleep/wake cannot produce a nonsense reading, and no worker timestamp is trusted. Reported as a median of five samples and withheld until a second sample exists: the first round trip after connect lands while the worker is still importing torch, which measured 139 ms on loopback and, averaged, carried that for a minute. This is CONNECTION latency, not time-to-result. It is shown as information, never as a routing input — RTT is milliseconds where inference is seconds, so ranking on it would optimise noise. Also fixes a bug the picker exposed: worker config was read from the pool, which caches the row handed to it at connect time. Renaming a CONNECTED worker updated the database and the API kept serving the old name until it reconnected — same for priority and enable/disable. Config now comes from the database and liveness from the pool, never the reverse, and writers refresh the live copy so the scheduler's logs do not use a stale name. Adds worker rename (the backend already supported it; no UI called it), worker address as seen by the control plane rather than self-reported, and ready/busy/offline status behind the header dot. |
||
|
|
43de1c794c |
feat(workers): remote GPU workers over a versioned gRPC protocol
Send individual jobs to GPUs on your other machines while everything else stays local. Opt-in, off by default: with the toggle off there is no listening socket, no certificate and no background loop. Design follows remote/goal_v2.md, the council-revised goal doc. The decisions that shaped the code, and why: * A disconnect is an unknown outcome, not a failure. The original design reassigned on disconnect while also describing the case where the worker had already finished — following both guarantees duplicate execution. An attempt now holds a grace window; a worker returning inside it commits its result and no second attempt is ever made. * At-least-once execution, exactly-once result commit. The result is persisted BEFORE it is acknowledged, so a crash between the two cannot silently lose a finished render. * Deadlines are phased (accept -> model load -> execute -> deliver) and liveness is a progress lease. The old fixed 30s execution budget was two orders of magnitude below what this product actually does; silence is the failure signal, not slowness. * Capacity is derived from free VRAM, never configured: a static value corrupts output under torch.compile thread affinity (#315) and aborts the process on small cards (#567). * A circuit breaker replaces the reliability-score/quarantine machinery, which had no recovery path (no probation workload exists in a TTS product) and penalised consumer networks for existing. * Identity is a keypair the worker generates and never sends. A server-assigned id is a name, not an authenticator, so revocation of one would be theatre. Enrollment tokens are single-use and carry the control plane's certificate fingerprint for pin-on-first-use. Adds the domain core, scheduler, durable task store, gRPC transport, worker agent, management API, Settings panel, and docs. Protobuf reserves the tenant/trace/usage fields a hosted control plane would need, since adding them later means upgrading a whole fleet. Includes tests for the failure paths that matter: duplicate delivery, stale-session fencing, reconnect reconciliation, grace expiry, breaker attribution, and a real end-to-end TLS round trip. |
||
|
|
f3286c5e6e |
feat(dub): paste a translation from an external source onto existing segments
After transcription the user can paste a translation produced elsewhere (ChatGPT, DeepL, a human translator) and have it map onto the segments that already exist — no re-transcription, no timing loss. Three input shapes are auto-detected: a timestamped .srt/.vtt (cues matched to segments by time overlap, greedy one-to-one so one long cue can't be copied onto several rows), numbered lines (`1.` / `2)` / `[3]`, mapped by number and falling back to order when a model renumbers mid-answer), and plain lines (positional, blank lines treated as separators rather than empty translations). Nothing is applied until the preview dialog has shown every row as before→after with unmatched rows flagged. Applying goes through `pasteTranslations` in useSegmentEditing, which mirrors `segmentEditField`'s duties across rows in ONE undo step: write `text` and `translations[dubLangCode]` in lock-step and clear the stale machine-translation badges. It never writes `text_original` (the translate source `handleTranslateAll` reads — overwriting it would poison every later re-translate) and never touches a language other than the active one. Changing `text` alone marks those rows stale via the existing per-language fingerprints, so no new flag is needed. The new `POST /dub/parse-subtitle-text` is a stateless wrapper over the existing `services.srt_parser.parse_srt`, so the lenient cue parsing stays single-sourced instead of being reimplemented in JavaScript. Also fixes a ReDoS in that parser, reachable today via /dub/import-srt: `_TIMING_RE` used `^\s*` under re.MULTILINE, so at every line start the engine consumed all remaining blank lines before failing on the first digit — quadratic. 20k blank lines already took 1.7s and a 2 MB blank-line file never returned, pinning the request thread. Horizontal-whitespace-only classes make the scan linear. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8111ace3ee |
test(api): add /system/last-run-crash + ack to the route snapshot (#1164)
Regenerated with scripts/dump_api_routes.py — the inventory guard (tests/test_api_route_inventory.py) rightly flagged the two routes the run-sentinel forensics added. |
||
|
|
018cdcb47f |
fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store (#1149)
* fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store Review round on #1148, both findings real: - Greptile P1 "missing translations leave mixed text": the in-browser translations map can be PARTIAL (tracks generated before per-language persistence, partial regens); the non-destructive switch then left those rows in the previous language under a single-language preview. New GET /dub/segments-text/{job}?lang= exposes segments_i18n (the authoritative per-language map every generate rebuilds); the tab click hydrates only the gap rows, failure-silent, and skips stale responses if the user switched again mid-fetch. - CodeRabbit "clear stale dialect": the dropdown paths each cleared a non-matching dubDialect by hand; the guard now lives inside switchDubLangCode so every caller (dropdown, multi-language loop, preview tabs, future ones) inherits it. Matching dialects survive. Tests: endpoint (i18n map served, never-generated track -> empty map, legacy job -> empty map), hydration (stored rows swap instantly, missing row hydrates from the mock backend and is cached into translations), dialect guard (cleared on mismatch, kept on match). Suites: dub sweep 262, frontend 1253, both green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(api): register /dub/segments-text in the route-inventory snapshot The inventory guard caught the new endpoint exactly as designed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a4d9d9f128 |
feat(dub): underrun fill — short dubbed lines are slowed toward their slot instead of leaving dead air (#1137)
* feat(dub): underrun fill — short dubbed lines are slowed toward their slot instead of leaving dead air The dub pipeline has always handled audio that is too LONG for its slot (atempo compression, Smart Fit's audio/video split, trims). Audio that is too SHORT was start-aligned and abandoned — and that is the common case, not the corner: translations routinely speak faster than the source delivery. Measured on a real 4-segment dub, 8.8 of 18.7 seconds of original speech time had no dubbed voice. What fills those holes is the separated bed's under-speech residue (37% of the original energy, measured), so the user hears them as BOTH "little silences" AND "the music is numbed" — and sees them as lip-sync failure, since the mouth keeps moving after the dub stopped. The fill: when a line's natural duration covers less than UNDERRUN_TOLERANCE (95%) of its slot, slow it toward the slot with the same pitch-preserving atempo pipe the compression path uses, bounded at min_audio_rate (default 0.85x — comfortably natural; atempo handles <1 natively). Wired into both fitting strategies: - fit_planner._fit_one: need < 1 now resolves to audio_rate=max(need, floor), status "audio_slowed" — planner stays a pure function; golden fixtures regenerated per their own instructions (10 substantive lines: five underrun segments across four scenarios flip to audio_slowed@0.85). - dub_generate smart_fit branch: applies the rate in both directions (the target formula was already direction-agnostic). - dub_generate strict_slot branch: mirror of its compression arm. - stretch_video and concise strategies deliberately untouched (natural-rate by design / never-intervene by design). OMNIVOICE_UNDERRUN_MIN_RATE overrides the floor (1.0 disables; clamped to atempo's sane range). The per-segment fit badge shows "slowed N.NNx" with a tooltip, translated in all 21 locales. Tests: planner contracts (fill bounded by floor, tolerance zone untouched, disable switch, empty-audio guard), the flipped unit/golden/integration expectations updated with the rationale, and the existing smart_fit integration test now exercises the fill through the real mix loop (its seg0 comes out audio_slowed@0.85 end to end). Full suite: 2987 backend + 1236 frontend. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(dub): strict-slot slow-downs report themselves honestly (review); ru pitch wording Review round on #1137: - Greptile P1 "slowdown reports fits" — REAL: the strict_slot underrun fill fell through to the unconditional {"status": "fits"} entry, so a slowed segment's badge hid the applied rate (and compression_applied mislabeled it). The branch now emits {"status": "audio_slowed", "audio_rate": …} like the smart_fit path — same honesty contract everywhere. - Greptile P1 "padded audio hides underruns" — REFUTED with evidence: nothing pads strict-slot audio before the check (_load_entry_wav returns the natural-length WAV; only error/silence slots are slot-sized, and those are synthetic silence by design). On-disk segment WAVs measure both shorter and longer than their slots, which pre-padding would make impossible. - CodeRabbit: Russian tooltip now says "высота тона сохранена" (pitch), not "высота сохранена" (height). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ea7a7b39b4 |
feat(privacy): opt-in analytics — hardened, off by default, enforced by code (#1120)
The rejected PR #1110 had a genuinely careful PII-free event design, but shipped three things a local-first app can't: exception autocapture ON (raw tracebacks — home paths, and in this codebase HF tokens out of exception messages — bypassing core.failure.sanitize() entirely), no user consent or disclosure, and 3,069 lines of PostHog wizard scaffolding. This is the same capability with those fixed. core/analytics.py, three rules, each enforced and tested rather than promised: 1. OFF unless the user says yes. TWO gates must both be true: a build-provided POSTHOG_PROJECT_TOKEN *and* the user's analytics_enabled pref, default False. A default install transmits nothing, so "nothing leaves your machine" stays literally true for everyone who doesn't opt in. A broken prefs file fails CLOSED. OMNIVOICE_ANALYTICS_DISABLED=1 is a hard kill switch above both. Withdrawing consent tears the client down immediately — no restart. 2. NO exception autocapture. Explicitly disabled; a test asserts the constructor arg, because the SDK's default is the leak. 3. Metadata ONLY, by allowlist. Every property passes sanitize_properties(), which DROPS any key not on _ALLOWED_PROPS and refuses long strings — so no future caller can leak a take's text, a path, or a voice name by adding a field. text_length is the LENGTH; the text itself has no way through. The person id is a random per-install UUID — not hardware, hostname, or username. UI: Settings → Privacy → "Help improve OmniVoice" states in the panel exactly what is sent, exactly what never is, and that it can be turned off — rather than burying it in a policy. No destination in the build (any source build) → the toggle isn't shown, because an inert switch would be a lie. Docs: README FAQ answers "does OmniVoice collect any data about me?" honestly. Also fixed a bug I'd introduced in my own wiring: the generation event referenced variables not in scope, and the call site's bare `except: pass` swallowed the NameError — so the event would have silently never fired. The call site now logs. 12 tests (default-off / opt-in without token still can't transmit / both gates / kill switch / consent withdrawal / prefs failure fails closed / allowlist drops text+paths+names / long strings refused / autocapture OFF / never raises / random install id). Backend 2936 passed; frontend 1211 passed. Refs #1110 Co-authored-by: mergetest <nizam4103@gmail.com> |
||
|
|
0e2c00a403 |
feat(privacy): Settings → Usage — local-only insights instead of cloud analytics (#1114)
* feat(privacy): Settings → Usage — local-only insights, the answer to cloud analytics A PostHog integration was proposed and rejected (PR #1110, closed): sending usage events to a third-party endpoint would break the one promise this product is built on — nothing leaves your machine — and local-first is the reason people choose it over ElevenLabs. But the question analytics was meant to answer ("how am I using this?") is a fair one, so answer it locally. services/local_stats.py aggregates the history the app has ALREADY written to the user's own SQLite DB: takes, audio produced, compute time, starred, active days, voices/dubs/projects/exports, and distributions by mode and language. GET /stats/usage serves it over loopback; Settings → Usage renders it. The three properties that stop this becoming telemetry by accident: - READ-ONLY. No new table, column, or event stream. Delete the feature and not one byte of stored data changes. - NO CONTENT. Counts and totals only — the `text` column of a take is never read and never returned; no paths, no ids, no person. Pinned by a test that asserts the payload contains no take text, no /Users/ path, no row id. - NO NETWORK. There is no client, no endpoint, no token. It has no way to send anything anywhere. The panel states the guarantee in the UI, because a privacy promise the user can't see isn't worth much. Route added to the API-surface snapshot (the inventory guard caught it, as designed — one line: GET /stats/usage). 4 backend tests (aggregation / never-leaks-content / empty install / missing table degrades to 0) + 4 frontend tests. Backend suite 2924 passed; lint, format, typecheck clean. Closes the analytics question opened by #1110. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(settings): use the real --chrome-fg-dim token in UsageTab (css-token guard) cssTokens.test.js is a frontend guard that every var(--…) a component references actually exists — an undefined custom property with no fallback is an invalid declaration, so the style silently does nothing. UsageTab referenced --chrome-fg-subtle, which doesn't exist; the dim sub-label token is --chrome-fg-dim (what the other settings panels use). My miss: I ran the full BACKEND suite but only the two new frontend test files, so this guard never ran locally. Full frontend suite now green (1211 passed). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bd85bab624 |
chore: retire finished planning archives from the repo root (#1095)
Removes ~110 files of process noise (all preserved in git history): .planning/ (GSD-era phases/quick-plans/issue-clusters; workflow retired 2026-07-08), specs/ (spec-kit specs for shipped features 001-007), design/ (pre-React ASCII mockups), research/ (legacy Gradio archive), and .agents/ (rules for a third-party agent tool no longer in use). The four load-bearing decision docs move to docs/adr/ with an archival note; every live pointer follows (gguf engine module docs + quant_map, inject-apprun.sh, pyproject/test comments, fixture README + its seed script — kept byte-identical). The CJK allowlist drops the deleted legacy_gradio entries; STRUCTURE.md and ROADMAP.md document the removal instead of linking into it. Backend suite: 2891 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3d8799d9d2 |
feat(asr): complete activation flow for the OpenAI-compatible remote ASR engine (#1087)
The openai-compat-asr backend (#877) shipped with settings routes but no discoverable activation path: the config panel hid in Settings → Models, its hint text claimed "there's no in-app engine picker for ASR yet" (stale — the matrix has one), and there was no way to check a server actually answers before pointing a dub/dictation run at it. Configure → test → activate now live on one screen, Settings → Engines: - The ASR family tab mounts the config panel (URL / model / optional API key) below the engine matrix; saving refetches the matrix via a new reloadToken prop so the engine row flips unavailable → available and its "Use" button appears without a manual refresh. - New "Test connection" button + loopback-gated POST /api/settings/asr-openai-compat/test: saves first (same stale-config contract as /llm-providers/{id}/test), then probes GET {base_url}/models — no audio leaves the machine. The structured verdict maps to localized, actionable messages: latency + whether the configured model is listed on success; classified auth_failed / http_error / timeout / unreachable / ok_no_models failures. detail is core.scrub-ed; the key is never logged or echoed. - Engine reads persisted config fresh per transcribe (regression test) — config changes need no backend restart. Never default-active: ASR auto-detect only picks local engines. - i18n for every new string (en.json); no hardcoded CJK; identical behavior on macOS/Windows/Linux (pure HTTP + React). - Docs-sync: docs/engines/openai-compatible-asr.md rewritten around the one-screen flow with LM Studio / llama.cpp / Groq / OpenAI examples and the privacy note; README engine table cell updated. Verified end-to-end against a fake OpenAI-compatible server: UI drive (configure → test → row flip → Use) plus a real transcription through the backend's /v1/audio/transcriptions immediately after a config change, no restart. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ff56865cf7 |
feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines (#1083)
* feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines
IndexTTS-2 required four manual terminal steps (git clone, uv venv,
uv pip install -e ., export OMNIVOICE_INDEXTTS_DIR). This turns that into
a guided in-app install:
- backend/services/sidecar_install.py — parametrized sidecar provisioner
(SidecarSpec/SPECS so future sidecar engines are one entry, not another
installer). Resumable background job with step-by-step status: disk-space
preflight (needs-X/have-Y message), source fetch (git clone --depth 1
primary, GitHub tarball fallback when git is absent/fails), dedicated
venv via uv (OMNIVOICE_BUNDLED_UV → PATH resolution; transformers<5
isolation preserved — the parent env is never touched), import-probe
verification, IndexTeam/IndexTTS-2 weights into <checkout>/checkpoints
(where the sidecar actually loads from) via snapshot_download with the
auto-selected/configured HF endpoint + token — no hardcoded
huggingface.co — and persistence of OMNIVOICE_INDEXTTS_DIR (os.environ
for immediate use, prefs.json env.* for the next launch). Idempotent:
partial installs repair, downloads resume, healthy installs (incl. a
user's own clone) report already_installed and are never touched.
- API: POST /engines/{id}/install starts the job, GET
/engines/{id}/install/status polls it, DELETE /engines/{id}/install
removes an app-managed install (loopback-gated; refuses user-managed
clones). list_backends() gains one_click_install.
- Frontend: Settings → Engines shows an Install button on the IndexTTS2
row with per-step progress, live log tail, weight-download %, and
error+remediation; the manual setup snippet is demoted to a collapsed
"Manual install" fallback. All strings via i18n (en.json).
- OMNIVOICE_INDEXTTS_DIR joins the Settings env-var allowlist
(single-sourced from the installer SPECS).
- Docs: docs/engines/indextts.md leads with the one-click flow; manual
steps become the fallback section. CHANGELOG Unreleased entry added.
- Tests: tests/test_sidecar_install.py (24 cases — happy path, disk-space
fail, git-absent/git-failing tarball fallback, partial-install repair,
already-installed/running gating, uninstall safety, spec↔bootstrap
contract, router wiring) + 6 new EngineCompatibilityMatrix RTL cases.
API route snapshot regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): harden the sidecar installer — review findings
- Route namespace: /engines/sidecar/{id}/install — a dynamic
/engines/{id}/install would shadow the literal
POST /engines/sonitranslate/install (engines router registers first);
regression-guarded by test_sidecar_routes_never_shadow_literal_engine_routes.
- Weights completion marker: a killed-mid-download multi-shard weights dir
(config.yaml + plausible shards) no longer passes for healthy; the marker
is written only after snapshot_download returns, so re-runs resume.
- _run_logged: drain thread + proc.wait(timeout) + POSIX process-group kill
— a grandchild holding the stdout pipe can no longer hang the step past
its timeout.
- Job log lock: the status poll's list(deque) copy no longer races the
worker's appends (RuntimeError under active logging).
- Self-heal: a healthy managed install whose env var was lost (prefs wiped)
is re-pointed by start_install instead of reported already_installed
while the engine stays unavailable; legacy bootstrap installs (Probe-2
venv) are trusted via the engine's own probe.
- Single-sourced uv/venv-layout resolution: engines.indextts.bootstrap now
delegates _locate_uv/_venv_python_path to services.sidecar_install.
- Frontend: stable poll interval (keyed on the running-id set, not the
status map), reload on a job that finishes before the first poll,
re-attach to an in-flight job on remount, i18n'd Install aria-label,
manual-install <details> auto-opens on failure, snippet block hoisted
out of the JSX IIFE.
- list_backends: sidecar-installable set hoisted out of the per-engine
loop; exhaustive-shape registry test updated for one_click_install.
- Tests rebind the live services.sidecar_install module per test (other
suites purge sys.modules["services"], which made router tests
order-dependent).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): fill in the PR ref (#1083)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): validated tarball fallback + scanner-clean installer
- The pre-filter= extractall fallback (Python < 3.11.4) now extracts
member-by-member behind the same guards extractall(filter="data")
enforces — regular files/dirs only, no absolute paths, no ../ escapes,
resolved-path containment. Kills the new CodeQL py/tarslip (high) and
Bandit B202 (error) alerts; regression-tested with a malicious tarball
(test_safe_extract_members_blocks_tar_slip).
- snapshot_download tracks the weights repo's default branch on purpose
(same policy as every other model download; artifacts are
checksum-verified by hf_hub) — documented + B615 waived at the call.
- Explanatory comments on the intentional empty-except blocks
(CodeQL py/empty-except notes).
Verified locally: bandit -ll -ii on the module reports 0 MEDIUM+ findings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): address Greptile review — Windows tree kill, prefs write race, poll robustness
- _kill_tree: Windows now uses taskkill /F /T so a git/uv helper spawned by
the timed-out child can't keep writing into the checkout (POSIX already
killed the process group). Unit-tested with os.name patched to nt.
- core/prefs: mutations (set_/delete) are serialized behind a module lock —
the installer worker persisting its env.* key concurrently with a Settings
write could previously drop whichever key saved first (whole-class fix:
every threaded prefs writer, not just the installer). Fail-before/
pass-after: tests/test_prefs_thread_safety.py.
- Matrix polling: at most one in-flight status request per engine (an old
'running' response can no longer land after a newer 'succeeded' and
restart the poller), and four consecutive poll failures drop the stale
snapshot instead of showing "Installing…" and hammering a dead backend
forever.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9c81e3389d |
feat(network): automatic Hugging Face endpoint selection — probe, pick, remember (#1082)
* feat(network): automatic Hugging Face endpoint selection — probe, pick, remember Restricted-network first-runs (the #984 class: huggingface.co unreachable, user dead-ends before discovering the mirror setting) now self-heal by default, while explicit endpoint choices are never second-guessed. - New backend/services/endpoint_race.py: parallel HTTPS reachability + latency probes of huggingface.co and the hf-mirror.com community mirror (3s timeouts). Probes are the only signal — no geo-IP, no third-party calls. Reachable beats unreachable; with both reachable the official endpoint wins unless the mirror is decisively faster (anti-flap hysteresis). The pick is cached in prefs and re-raced only on first run, a network-classified download failure, staleness (>7 days), or an explicit "Test again". - Manual mode is sacred: HF_ENDPOINT env, an hf_endpoint pref, or any explicit Settings pick disables auto-switching entirely; OMNIVOICE_HF_ENDPOINT_MODE=manual is a hard opt-out. - Wiring: the wizard preflight races endpoints when nothing is configured (honest copy when the mirror wins; warn-not-block when nothing is reachable); Model Store installs and the model-cache auto-repair resolve their per-call endpoint= through the cached decision, and a network-classified failure re-races once per repo per process and retries on the new winner (same guard pattern as the cache-recovery ladder). - Settings → Models → Hugging Face mirror gains "Auto (recommended)": shows the current pick, measured latency, last-checked time, and a "Test again" button (POST /api/settings/hf-mirror/test). Existing explicit configs surface as the matching manual mode. Panel notes that hf_hub checksums every download regardless of endpoint. - Tests: policy/cache/failover matrices in tests/test_endpoint_race.py, preflight + settings + repair-failover integration with mocked probers, HFMirrorPanel mode tests, and a suite-wide conftest guard that pins the probers so no test can hit the real network. - Docs: downloading-models.md and install/troubleshooting.md describe the automatic default and both opt-outs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): Unreleased entry for automatic HF endpoint selection Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint-probe pin uses an isolated MonkeyPatch and clears the decision cache; dtype guard tolerates stubbed torch The autouse probe pin requested the shared monkeypatch fixture, hoisting its setup earlier for every test and reordering teardown against the fp16 guard — which then ran torch.get_default_dtype() on test_torch_compile_gate's SimpleNamespace stub. The pin now uses its own MonkeyPatch context and also clears the prefs-cached endpoint decision per test (one test's auto pick leaked into other tests' preflight labels on CI ordering). The dtype guard additionally skips non-module torch stubs outright. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(tests): endpoint env vars can no longer leak out of the mirror-settings suite set_hf_mirror writes os.environ[HF_ENDPOINT] during the test, and monkeypatch.delenv(raising=False) on an absent var records nothing to undo — so the write leaked process-wide and flipped later suites' preflight network checks into the explicit-endpoint branch (the CI-order failures). Guaranteed save/restore autouse fixture at the source, plus defensive env shedding in the preflight suite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5562aa16a7 |
feat(setup): media tools become invisible — bundled by default, controllable in Settings (#1071)
Most users should never learn what ffmpeg is. The Setup Wizard's SYSTEM
PREFLIGHT stops listing FFmpeg / FFprobe / yt-dlp as user-installed
requirements ("brew install ffmpeg…"): they are internal dependencies the
app provisions for itself. Genuine user facts (OS, RAM, disk, GPU,
network, Python) are untouched.
Backend
- New services/media_tools.py: per-tool status {version, path, origin:
sidecar|bundled|system|custom}; background acquisition of a pinned,
SHA-256-verified static ffmpeg+ffprobe build (immutable-commit fetch
from the same upstream the static-ffmpeg pip package uses — that
package itself was audited and rejected: mutable raw/main URL, no
checksums, writes into site-packages); binaries are `-version`-probed
via the existing _binary_runs before being trusted, installed under
DATA_DIR (update-surviving, frozen-build-safe), zero new Python deps.
- ffmpeg_utils resolution chain gains the acquired-bundled tier — and
ffprobe finally has a bundled tier at all (imageio-ffmpeg ships none),
closing the source-install gap.
- New /media-tools router (loopback-gated, same contract as
/system/set-env): status, acquire, {tool}/custom-path | use-system |
restore, ytdlp/update | restore. Overrides persist via the existing
env.FFMPEG_PATH / env.FFPROBE_PATH prefs convention — one store, no
competing controls.
- yt-dlp updates: audited in-venv pip/uv upgrade and rejected (venv is
uv-managed with no pip; yt-dlp is a locked dep, so the updater's
--inexact drift sync would revert it). Instead the newest wheel —
verified against PyPI's own sha256 — lands in a DATA_DIR overlay
prepended to sys.path at startup: survives app updates, works in
frozen builds, and "Restore tested version" is just deleting the
overlay. Gallery now runs yt-dlp via `python -m yt_dlp` (module, not
PATH) so the CLI can never be a user-install task either.
- /setup/preflight drops the three tool rows, carries a media_tools
verdict, and self-heals: kicks the bundled download in the background
when no tier resolves (never re-fires after a failure — the wizard's
card owns Retry). diagnose + the ffmpeg-missing notification now point
at Settings → Audio tools instead of package managers.
Frontend
- Wizard: new MediaEngineCard — renders NOTHING when the engine is ready,
a one-line progress while acquiring, and only on failure an actionable
card (Retry / Use a system copy / Choose file…).
- Settings → Audio tools (new category, System group): FFmpeg + FFprobe
rows with version, path, origin badge, Use system copy / Choose file… /
Restore bundled, header-level "Update bundled build"; yt-dlp row with
Update + Restore tested version (+ restart affordance). Package-manager
commands appear only as copyable prose, never executed.
- The FFmpeg-path override moved out of Settings → Network (pointer row
deep-links to Audio tools; no second writer of env.FFMPEG_PATH).
Notifications gain a settings-tab action type.
- All strings i18n (en + defaultValue), a11y labels on every control.
Tests: 29 new backend (origin classification, checksum/size/probe
rejection, override persistence, overlay update/restore, router gating +
route-shadowing) + preflight contract tests (tool rows gone, verdict
present, auto-acquire fires once); 14 new frontend (wizard hide/progress/
failure-card, Audio tools rows/badges/actions). Route snapshot
regenerated. Docs (macos/linux install, troubleshooting §7b) describe the
new reality in the same commit.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
b603b9f78d |
fix(settings): factory reset covers all prefs, guarded log clearing, temp reclaim, full Performance i18n (#1061)
Settings system-group cleanup — every fix keeps existing behavior contracts and adds a fail-before/pass-after regression test: - Factory reset now does what it promises: clears every locally-persisted preference via a single registry (utils/prefKeys.js) instead of only the zustand blob — nav-rail side, capture live-typing, stories speed, logs footer state, last settings category, dismissed tips, donate prompts, and the legacy omni_ui blob included. User data and connection state (omni_transcriptions, ov_backend_url, ov_api_key) are explicitly preserved, and prefKeys.test.js scans the source tree so any future localStorage key must be categorized or CI fails. The failure toast now carries the actual error message. - Disk-usage "Clear logs" is confirm-gated with the same wording as Settings → Logs — it truncates the crash log (the bug-report artifact), so it can no longer be a single stray click. - Temporary files got a reclaim action: a confirmed "Clear temp files" button backed by POST /api/settings/storage/temp/clear, which deletes only the omnivoice* entries in the OS temp dir (symlinks unlinked, never followed) and invalidates the cached report. - Performance panel goes through i18n end to end (title, row, note, hint, errors, aria-label) — it was the last fully hardcoded panel; the non-Windows subtitle now reads "Windows only — not needed on this platform" instead of "not applicable". - History retention: GET failures now surface an alert and hold Save until a load succeeds (404 from older backends stays silent), Enter saves, the dead !res.ok branch is gone, and the bespoke button is the shared Button. - Logs tab: "Open folder" reveals the log file, "Copy visible log" copies the tail, the viewer autoscrolls to the newest lines, and the scroll box is keyboard-focusable (role=log) with a labelled source switcher. - Storage paths: the app-data row is labelled "App data stored at" (it was borrowing the Privacy tab's "Uploads stored at"), and all three path rows gained Open folder. - i18n stragglers routed through t(): storage load/open/clear fallbacks, the backend-status badge, and the frontend log buffer label. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
efc99be337 |
feat(studio): generation takes — star, replay, and restore past takes; capped history retention (#1052)
Every generate already recorded a generation_history row; now that history is usable: a takes rail in the workspace history lists recent takes with star/ unstar, replay, and one-click restore as the active output. Alembic migration 0009 adds the starred column (the startup schema self-heal covers pre- migration DBs), a retention cap (setting, default 200) prunes the oldest UNstarred rows — starred takes are never pruned — and history WAVs are only deleted when no other row references them. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5a7d9cc05c |
feat(asr): generic OpenAI-compatible transcription backend (#877) (#1003)
First slice of the community's two-track proposal for #877: a generic OpenAI-compatible ASR backend that works TODAY, without waiting on transformers to ship a direct Qwen3-ASR integration (tracked separately, still blocked upstream). Points OmniVoice's transcription at any server exposing POST /v1/audio/transcriptions — a self-hosted Qwen3-ASR/ FunASR/SenseVoice server, or OpenAI's own API. - New OpenAICompatASRBackend (backend/services/asr_backend.py): a pure network client, no local model, no install. Prefers response_format=verbose_json for real per-segment timestamps, degrades to plain text (matching MoonshineASRBackend's shape) when a minimal server rejects that format. Never leaks a raw SDK/httpx exception to the caller (#977 convention) — wraps network/auth failures in a clean, actionable RuntimeError naming the server. - Settings persist via the same encrypted-secret convention as services/llm_providers.py (settings_store.set_secret for the API key — Fernet-encrypted, never a .env row, never echoed back; get_text/ set_text for base_url/model). New GET/PUT /api/settings/ asr-openai-compat, loopback-gated like every other settings route. - Frontend: a small settings panel (Settings → Models) mirroring HFMirrorPanel's exact structure. No ASR engine picker exists yet for ANY ASR backend (only TTS has one) — activating this engine still needs OMNIVOICE_ASR_BACKEND=openai-compat-asr; documented plainly rather than pretending otherwise. - README's ASR Engines table (9 → 10 engines) and docs/features.yaml's drift-checker inventory updated; the '9 engines, all fully local' claim corrected since this one genuinely isn't. - docs/engines/openai-compatible-asr.md: setup steps + an explicit privacy note (unlike every other ASR engine, audio leaves the machine to whatever server is configured). Regression tests: tests/test_asr_openai_compat_877.py (12 tests) — is_available() gating, verbose_json + plain-text response adaptation, network-failure error hygiene, SDK retry disabling, and the settings endpoints' persist/mask/clear-vs-unchanged semantics. Fixed two real full-suite-only failures found during verification (not brushed aside): the API route inventory snapshot needed regenerating for the two new routes, and this file's own tests had a module- staleness bug — a collection-time settings_store import went stale relative to a test-time-fresh fixture when another test elsewhere in the ~2400-test suite reimports the module — fixed by making settings_store itself a fixture resolved at test-run time, same lesson already applied to tests/test_mm2_lifecycle.py earlier this session. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5bd8968aea |
feat(engines): real synthesis "Self-test" + copy-paste setup snippet for opt-in engines (#930)
* feat(engines): real synthesis "Self-test" + copy-paste setup snippet for opt-in engines Builds on #905's Engines-settings fixes (verified still green: license dialog mounts, matrix reloads on select, cpu_fallback routing toast, cpu-native → cpu_only). Two enhancements, no #905 behavior touched. Real "Self-test" for in-process TTS engines ------------------------------------------- The existing /engines/{id}/health probe only imports the package and reports "deps OK" for in-process engines — it never proves the engine can emit audio. New POST /engines/{id}/selftest runs a *tiny real synthesis* from a fixed short ASCII phrase and reports ok + duration + sample-rate + sample count, proving the engine actually produces audio. Guardrails keep it cross-platform-identical and CPU-cheap: TTS + available + in-process only, bounded wall-clock timeout (OMNIVOICE_SELFTEST_TIMEOUT_S, default 90s) that returns ok=false/timed_out instead of hanging the panel, a process-wide lock so a click-storm can't stack model loads, loopback-gated, and only ever on user click (never on load). The Compat Matrix gains a "Self-test" button (with cooldown) that renders "0.82s @ 24 kHz in 820 ms". HF tokens in a synth error are redacted like the health route. Verified end-to-end: kittentts synthesized 89,200 samples @ 24 kHz. Copy-paste setup snippet for path-gated opt-in engines ------------------------------------------------------ IndexTTS / MOSS-v1.5 / dots.tts / Confucius4 gate on an OMNIVOICE_*_DIR env var. list_backends() now emits a single-sourced `setup_snippet` (the exact `export VAR=/path/...` line) surfaced with a Copy button inside the matrix's "Why unavailable?" disclosure, so users don't reconstruct it from the docs. Also tightened the incomplete SelectEngineResponse TS type to include the routing echo (routing_status/effective_device/routing_reason) the post-select toast already reads at runtime. Tests: backend selftest success/subprocess-reject/unavailable/unknown/loopback/ exception-capture/timeout/HF-redaction + setup_snippet shape; frontend self-test render, timeout marker, subprocess+ASR gating, setup-snippet render. New route added to the API route snapshot. Full vitest (808) + backend engine/routing/asr/ route-inventory/no-CJK green; lint 0 errors; format + typecheck:ci clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): allow setup_snippet key in list_backends shape assertion The engine self-test PR added setup_snippet to each backend entry but only updated the route-shape test; test_list_backends_shape strict-asserts the key set. Add setup_snippet there too. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8f71c90f20 |
feat(settings): LLM Skills — per-feature enable/route control for every LLM call (#912)
New Settings → System → LLM Skills area: every LLM-powered capability
(Cinematic & Autofit translation, speech-rate slot fitting, glossary
auto-extract, direction parsing, dictation cleanup) becomes a "skill" the
user can toggle or route to a specific provider (local Ollama/LM Studio vs
a remote key) instead of everything riding the one global active provider.
Backend:
- services/llm_skills.py — skill registry + settings_store persistence
(llm_skill.<id>.enabled / .provider), resolution precedence
override > active > none, resolve_skill_client() (OpenAI-compat client
bound to the effective provider; None when disabled/unconfigured) and
skill_backend() (OffBackend when disabled — the exact no-LLM object every
caller already degrades on).
- All five consumption points wired through the registry; a disabled skill
degrades exactly like "no LLM configured" today (Fast translation
fallback, refinement pass-through, heuristic direction parse, no-llm slot
fit, 503 on glossary auto-extract). No new degradation modes; defaults
(enabled + no override) keep existing setups byte-identical.
- OpenAICompatBackend gains an optional bound provider (None = active, the
historical behavior).
- GET /api/settings/llm-skills + PUT /api/settings/llm-skills/{skill_id}
(404 unknown skill/provider); route snapshot updated.
Frontend:
- LLMSkillsPanel (Sparkles, next to LLM Providers): one row per skill —
i18n name/description, enable toggle, provider Select ("Use active
provider" + configured providers, local ones tagged), ready /
needs-setup badge linking to LLM Providers. All strings via t()
(settings.llmskills_*).
Tests: 30 backend (precedence, per-consumption-point disabled semantics,
endpoint round-trips, validation) + 4 panel render/PUT tests. Docs:
translation-engines.md gains an LLM Skills section.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
16294fed44 |
feat(updates): data-safe updates — pre-migration DB backups, guarded venv heal, release notes + changelog reader (#909)
Backend: - core/db_backup.py: WAL-safe SQLite snapshot to omnivoice.db.backup-<version>-<n> before pending alembic migrations run; keep newest 3, prune older; skip >500MB with a log line. Restore is never automatic. - core/db.py: _run_alembic_upgrade now plans the run (up_to_date / pending / unknown_revision), snapshots first when migrations will execute, and raises MigrationError on a mid-flight failure — startup stops with the backup path named instead of continuing on a half-migrated DB. The #552/#547 unknown-revision class stays non-fatal (warn + additive reconcile). - core/changelog.py + GET /api/settings/changelog: parse the shipped CHANGELOG.md (single-line and wrapped bullet styles) into structured releases. - GET /api/settings/db-backup: newest pre-migration backup for the panel. Rust (bootstrap.rs): - #314 heal guard: an exit-signature match alone can no longer delete the venv — venv_rebuild_justified requires a structural problem or a failed direct interpreter probe; a venv that probes healthy is kept and the real error surfaced. Drift/repair remains in-place `uv sync` (non-destructive). - CHANGELOG.md now ships as a bundle resource and is copied/refreshed into the project dir so the changelog endpoint works in packaged installs. Frontend (Settings → Updates): - Available update shows its actual release notes (updater metadata body) through a safe markdown-lite renderer (text nodes only, refs stay plain). - "Your data is backed up before every update" line with the latest backup timestamp from the new endpoint. - "What's new" changelog reader (accordion, newest expanded) over the shipped CHANGELOG.md; GitHub releases list reuses the same renderer. - One-time, non-blocking "What's new" footer pill after an update (persisted last-seen version; fresh installs baseline silently). - All strings via t() with en keys (other locales fall back to English). Tests: db backup/rotation/failure-path units, migration-safety units, changelog parser (both bullet styles + real CHANGELOG.md), endpoint tests, route inventory regenerated, Rust decision-logic + probe tests, vitest suites for renderer/viewer/panel/pill logic. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b3c18db33f |
feat(settings): Storage panel — real disk usage, category breakdown, and low-space warnings (#906)
Settings → Storage now opens with a Disk usage panel backed by a new loopback-gated GET /api/settings/storage endpoint: - Per-volume totals (grouped by st_dev) + du-style sizes for everything the app owns: the HF model cache (with its ~10 largest models), the app data dir broken into voices/outputs/dub_jobs/batch/preview/ database/logs/other subtotals, engine venvs (backend/engines/*/.venv + the app venv), and omnivoice* entries in the OS temp dir. - Bounded scanning: per-category 10 s deadline → partial totals with an "unreadable" warning instead of a hung request; results cached in-process for 5 minutes, ?refresh=1 forces a rescan; the walk runs in a worker thread so the event loop never blocks. - Server-side warnings reuse the setup wizard's MIN_FREE_GB: free < min → critical, free < 2×min → low, volume holding the cache/data >90% full → volume_pressure, unreadable/timed-out paths → unreadable. The panel renders severity-colored banners, a data-volume gauge, proportion bars per category, Open-folder buttons (existing /export/reveal pattern), a Model Store jump for reclaiming model space, and the existing clear-logs action on the logs row. A critical warning is also surfaced outside Settings via the app-wide toast — once per session. All strings via i18n (en fallback). Tests: tests/test_storage_report.py (sizes, thresholds, cache/refresh, timeout partials, endpoint wiring) + StorageUsagePanel.test.jsx (categories, banners, once-per-session toast, refresh=1, error state); route added to tests/fixtures/api_routes.txt via the dump script. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
da9315815d |
feat(settings): LLM provider testing pass — latency + classified errors, model discovery, full i18n, router tests (#887)
Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
29269b9cf0 |
feat: LLM Providers page + Autofit translation quality (fit-to-segment-time) (#838) (#854)
* feat(llm): multi-provider LLM registry + encrypted key storage + settings API (v0.3.8, phase 1)
Foundation for the LLM Providers settings page and timing-aware (Autofit)
translation. Every provider in the shipped .env is OpenAI-compatible, so one
client drives all of them via a registry instead of a class-per-provider.
- llm_providers.py: registry of 16 providers (OpenAI, OpenRouter, Groq,
Cerebras, Google AI, Mistral, Cohere, NVIDIA, GitHub Models, Cloudflare,
HuggingFace, SambaNova, SiliconFlow, + local Ollama/LM Studio + Custom).
Field resolution precedence env → encrypted store → default; active-provider
selection (LLM_DEFAULT_PROVIDER → stored → first keyed remote; local requires
explicit pick so we never assume a local server is up). Legacy TRANSLATE_*
maps to the Custom provider (keyless-with-base_url preserved).
- settings_store.py: generic ENCRYPTED secrets (get/set/clear_secret,
list_secret_names) reusing the HF-token Fernet path; get_text/set_text now
refuse the secret namespace (no ciphertext leak).
- llm_backend.py: OpenAICompatBackend resolves the active provider's
base_url/key/model from the registry. Backward-compatible.
- settings API: GET /llm-providers, PUT /llm-providers/{id} (encrypted key +
overrides), POST /llm-providers/active, POST /llm-providers/{id}/test.
Loopback-gated; never returns key material.
- 11 registry tests; existing llm-endpoint/openai-available tests still green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(settings): LLM Providers page — configure any provider's key/URL/model + Test + set active (v0.3.8, phase 2)
New Settings → System → LLM Providers pane (Brain icon, searchable). Lists all
16 registry providers; pick one to configure its encrypted API key, base URL,
model (and Cloudflare account id), Test the connection with one round-trip, and
'Save & use for translation' to make it the active provider for Cinematic/
Autofit. Keys are write-only from the UI (masked placeholder, never echoed);
env-set keys show as read-only. Local providers (Ollama/LM Studio) need no key.
- LLMProvidersPanel.jsx: provider selector + per-provider config + Test/activate,
following the LLMEndpointPanel pattern (apiJson/apiFetch/apiPost, SettingsSection
primitives).
- settingsCategories.jsx: new 'llm-providers' category under System + Brain icon.
- Settings.jsx: route the category to the panel.
- en.json: settings.llm_providers label.
- Frontend build passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(translate): Autofit quality style + one-click LLM setup from the dub menu (v0.3.8, phases 3-4)
Autofit = Cinematic + a strict 'never exceed the segment time' fit. The LLM
rewrites each translated line so its target-language reading time fits within
the slot, preserving the video timing without harsh audio time-stretch.
Backend:
- speech_rate.adjust_for_slot(strict=): strict caps the accepted upper ratio at
1.0 (fit within slot) vs Cinematic's 1.08; best-effort, degrades gracefully
with no LLM.
- dub_translate: quality='autofit' takes the LLM refine path and runs the fit
pass with strict=True; reports quality_used accurately.
- TranslateRequest.quality doc note.
Frontend:
- 'autofit' added to the quality control (Settings Translation + dub menu) and
the TranslateQuality type.
- Dub menu: picking Cinematic/Autofit with no LLM no longer dead-ends on a toast
— it offers a one-click 'Set up' that routes to Settings → LLM Providers,
with copy about fitting translations to segment time (#838).
- 4 strict-fit tests; frontend build green; i18n keys added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(translate): document Autofit quality + the LLM Providers page (v0.3.8, phase 5)
- CHANGELOG [0.3.8] Added: Autofit style + LLM Providers page.
- docs/dubbing/translation-engines.md: Fast/Autofit/Cinematic quality section
and an LLM Providers setup section (16 providers, encrypted keys, offline
Ollama/LM Studio, env overrides).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(api): add /api/settings/llm-providers routes to the route-inventory snapshot
Regenerated tests/fixtures/api_routes.txt for the 4 new LLM-providers endpoints
so test_route_inventory_matches_snapshot passes (keep-main-green).
* style(frontend): oxfmt the LLM Providers panel + dub quality control (format:check green)
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ee638bda6c |
feat(tts): user pronunciation dictionary (expressive-tts slice 1) (#685)
* feat(tts): user pronunciation dictionary (expressive-tts slice 1) Per-term, per-language pronunciation overrides applied to text before synthesis, so names, brands, and acronyms come out right across generate, longform, and dub. Closes part of the #1 perceived-quality gap vs ElevenLabs (pronunciation dictionaries). First slice of docs/specs/01-expressive-tts.md. - Schema: additive `pronunciation_entries` table (alembic 0008, mirrored into _BASE_SCHEMA; tested upgrade — idempotent, downgrade, converge, back-compat). - Service: extend pronunciation.py to load enabled entries (cached) and apply longest-first, word-boundary-aware, per-language (global '*' + lang match, lang overrides global), reusing the existing ReDoS-safe matcher. - Inline one-off `[[term|replacement]]` overrides that don't persist and don't collide with [voice:]/[pause]/[Name]/SSML-lite (resolved pre-chunking). - API: /pronunciation CRUD + /test dry-run + import/export (loopback-guarded). - Apply point: generation.py after language resolves, before chunking — covers native + pluggable engines. - UI: PronunciationPanel in Settings → General; all strings via i18n. - Tests: migration lifecycle, CRUD, per-language, precedence, inline override, apply-at-synth. Route snapshot regenerated (+7). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(security): bound inline-override regex (ReDoS) + annotate parameterized UPDATE CodeQL flagged py/polynomial-redos on the [[...]] inline-override regex: [^\]] also matches [, so an unterminated run of [ allowed O(n) rescans from O(n) positions. Bound the inner class to {0,256} (linear; an inline override is a short respelling). Annotate the dynamic UPDATE (B608) — its column fragments are fixed literals and every value is a bound parameter; not an injection vector. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
022a3bd6b9 |
feat(dictation): live local dictation via sherpa-onnx + Voice settings panel (#683)
* feat(dictation): live local dictation via sherpa-onnx + Voice settings panel Add a sherpa-onnx ASR engine alongside the existing Whisper/NeMo dictation path, powering a genuinely live experience: as you speak, words type straight into the focused field (streaming partials via a new simulate_type command, self-correcting with backspaces) and commit per pause. Backend: - SherpaDictationBackend + sherpa_dictation registry of the 7 models (Parakeet TDT v3/v2, streaming Zipformer EN/ZH/bilingual, Paraformer bilingual, Whisper Tiny) from csukuangfj/* int8 HF repos; CPU provider for cross-platform parity. - /dictation/models + /dictation/prefs router; get_capture_asr_backend() honors the selected dictation model. get_active_asr_backend() (dub transcription) and the legacy WebM/Opus capture path are untouched. - True streaming over /ws/transcribe (OnlineRecognizer: live partials + per-endpoint finals); offline models surface partials via short re-decode. Frontend: - New "Voice" settings panel (enable, Toggle/Hold mode, model picker with offline/streaming/recommended badges + per-model download/delete). - Live word-by-word typing via simulate_type (enigo) with prefix-diff delta and backspace correction; paste fallback retained, no double-insertion. Deps: sherpa-onnx>=1.13.3 (+ sherpa-onnx-core); uv.lock regenerated, Docker frozen-install verified. API route-inventory snapshot updated. 40+ new tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dictation): register sherpa-onnx-asr engine in README + features inventory Fixes the docs-drift CI guard: the new sherpa-onnx-asr ASR engine existed in the registry but not in docs/features.yaml or README. Adds the live-dictation engine row to the ASR Engines table, bumps the engine counts (8→9), and adds the inventory entry. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): fold live-dictation into the [0.3.8] section main is 0.3.8 (untagged), so the dictation feature belongs in that release, not a separate [Unreleased] block. Merge the two Added lists under one [0.3.8], refresh the headline to lead with live dictation, and correct the capture description to reflect live word-by-word typing (not paste-on-pause). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
de80856cd9 |
test: backend route-inventory + webUI feature-coverage guards (#609)
* test: backend route-inventory snapshot + webUI feature-coverage guards A reusable testing system that verifies every feature surface is present: - tests/test_api_route_inventory.py: boots the app, diffs all 213 routes vs a committed snapshot (tests/fixtures/api_routes.txt), guards a critical-endpoint set, and floors the route count — any endpoint drift fails CI. - scripts/dump_api_routes.py: regenerates the snapshot. - frontend featureCoverage.test.js: every AppMode has a render branch, every lazy-imported page file exists, every feature has an i18n namespace. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): note the feature-coverage test system * test(api-inventory): isolate via subprocess + exclude env-dependent mounts CI surfaced two flaws in the first cut: - the in-process app import + sys.modules purge polluted later DB-touching tests (a cascade of 404s in test_dub_subtitles_309 etc.); - the snapshot included StaticFiles mounts (/demo_audio) and a conditional GET / root that register based on filesystem state, so a macOS-generated snapshot didn't match a fresh Linux CI runner. Compute routes in an isolated subprocess (scripts/dump_api_routes.py --print) and cover only the deterministic router surface (drop Mounts + root). 209 routes; inventory + previously-polluted tests now pass together. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
276875c397 |
feat(longform): canonical Python parser + golden corpus (#27 slice A) (#465)
The longform marker dialect (# heading / [voice:] / [pause] / SSML-lite) was parsed by three independent code paths that already disagreed (client vs server on [pause] units, [voice:] empty, H1-only chapters). This lands the single canonical Python parser; the JS port + cross-impl test follow in slice B. - New backend/services/longform_parser.py — parse_script_to_spans(text, *, default_voice, default_speed) + _parse_chapter_body (the reusable voice→pause →SSML layering the JS twin mirrors). Moves the H1/voice regexes verbatim from audiobook.py (already CodeQL-cleared), reuses parse_pause_markers + ssml_lite unchanged. Coerces None→"" and normalizes CRLF/CR→LF at entry (cross-platform parity so Windows-authored scripts never carry a stray \r). Adds default_speed plumbing (inline SSML speed overrides the per-line default). - audiobook.py: parse_audiobook_script is now a thin wrapper that wraps the canonical span dicts in Span/Chapter/AudiobookPlan — public return type and .to_dict() shape unchanged, all four router call sites untouched. Deleted _parse_spans / _HEADING_RE / _VOICE_RE and the now-dead `import re` + parse_pause_markers import. - tests/fixtures/longform_parser_cases.json — 78-case golden corpus (≥40 required) covering §A–I: H1-only chapters (H2–H6 + `# ` no-title → body), the full pause dialect incl. the NO-MATCH boundary, banker's-rounding ties ([pause 0.5]→0, [pause 1.5]→2), [voice:] empty→default, [voice:[nested]] literal, SSML nesting/spell/unknown-tag, speed override, CRLF, combined precedence. Generated from actual parser output (the truth the JS port must match). - tests/test_longform_parser.py — parametrized over the corpus + None-input + ReDoS-linearity (5000× repeats < 1 s). 130 passed (corpus + test_audiobook + test_pause_markers + test_ssml_lite all green); CJK guard green. |
||
|
|
9162f2b9e7 |
feat(stream): sentence-by-sentence /ws/tts via ported chunker (Wave 1.4) (#358)
Ports Patter's SentenceChunker (MIT, attribution header) behavior-identical — all 61 upstream golden parity scenarios ship as fixtures and pass, including documented quirks (current_behavior xfail semantics mirrored from their parity runner). Terminator tables carry functional CJK; file added to the test_no_hardcoded_cjk allowlist per convention. /ws/tts now splits the request into sentences and synthesizes each in turn, streaming the first sentence's PCM while later sentences are still generating — the time-to-first-audio win on multi-sentence input. Single-sentence requests behave exactly like the old single-shot path; 'start' metadata still waits for the first generation so lazy-loading engines report their true sample rate. Italian comma-decimal guard hard-disables aggressive first-clause flush per upstream. Spec 8a (docs/competitive-analysis.md) / parity program Wave 1.4. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4b21f82619 |
feat(dub): Smart Fit timing strategy — planner, fingerprints, generate path (phase A) (#347)
* feat(dub): Smart Fit planner, fit fingerprints, shared ffmpeg stretch helpers - services/fit_planner.py: pure, I/O-free planner for dub-length fitting v2 — slack absorption (gap guard), audio-only band (<=1.2x), geometric 50/50 audio/video split capped at 1.5x / 2.0x, residual overflow accounting, and a stretch_video-compatible video_plan + fitted timeline cursor. Clean-room reimplementation from a published description. - services/incremental.py: fit_fingerprint() over the fit params with the same _canon_value canonicalisation as segment hashes (#281 class). Fit params stay OUT of segment_fingerprint — a fit change re-mixes, never re-TTSes. - services/ffmpeg_utils.py: move _atempo_chain/_pitch_preserving_stretch out of the dub_generate router (lazy torch/numpy imports) so the Phase B export pipeline can reuse them; add probe_duration() ffprobe helper. - schemas/requests.py: timing_strategy gains "smart_fit"; optional fit_options knob overrides default server-side. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(dub): smart_fit branch in the generate path TTS loop unchanged (dur_s=None, natural-rate WAVs on disk). After the loop, plan_fit() decides per segment; the mix loop applies audio_rate via the pitch-preserving atempo pipe (linear-interp fallback), trims residual overflow with the existing fades, and places audio at the planned new_start on a fitted-length canvas. Truthful fit_status entries (audio_rate / video_ratio / overflow_s) feed the row badges. Persists job["fit_plans"][lang] = {plan (exact _build_video_stretch_filter_graph shape), fitted_segments (cue times from ACTUAL stretched sample positions), total/orig duration, params, fit_fp} and mirrors fit_fp on dubbed_tracks[lang]. video_stretch_plans untouched. Strategy-transition guard: job["seg_wav_kind"] records whether on-disk seg WAVs are natural or slot-squeezed; a smart_fit partial regen over slotted (or unknown) WAVs forces one full regen instead of double-compressing. Old strategies and old persisted jobs are byte-identical (all new reads via .get()). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ui): Smart Fit option in the dub timing picker (all 21 locales) - prefsSlice: TimingStrategy union gains 'smart_fit'; optional FitOptions overrides (null by default — backend defaults apply identically on every platform); persisted alongside timingStrategy. - DubTab: Segmented gains Smart Fit with i18n label + tooltip. - useDubWorkflow: sends fit_options only when set and strategy is smart_fit. Default strategy stays 'concise' — no default behaviour change on any platform. - locales: dub.timing_smart_fit{,_title} translated in all 21 languages. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(dub): fit planner unit + golden suites, smart_fit generate-path integration - test_fit_planner.py: threshold boundaries (0.9/1.0/1.2/1.21/4.0), cap saturation -> overflow, slack absorption incl. gap guard, last-segment tail, cursor monotonicity, allow_video_retime=False, video_plan fed straight into _build_video_stretch_filter_graph, fit_fingerprint canonicalisation (int vs float, omitted vs default — the #281 class) and a pinned stable digest. - tests/fixtures/fit_planner/*.json: 4 golden FitPlans; algorithm drift is a deliberate fixture diff, never a silent change. - test_smart_fit_generate.py: hermetic end-to-end runs (mock TTS, no ffmpeg) covering audio-only stretch, hybrid timeline growth + persisted plan shape, fit_options override, strict_slot->smart_fit forced regen then zero-TTS fit-only re-mix, and concise back-compat. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(competitive): dub-length fitting row reflects Smart Fit Phase A Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(incremental): mark fingerprint hashes usedforsecurity=False — dedup keys, not security (Bandit) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0fbc65f29d |
test: scrub brand name from whisper segmentation fixture (#162)
The whisper_screenshot transcription fixture + its segmentation test referenced a real product/brand name. Swap it for the neutral placeholder 'Acme' (fixture text + chunks + the expected-segment assertions), keeping the test's consolidation behavior identical. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c3695e1668 |
Phase 2 Plan 02-03: IndexTTS on SubprocessBackend (closes #42) (#98)
Migrates IndexTTS-2 off the in-process import path and onto the SubprocessBackend primitive shipped in Plan 02-01. Closes issue #42 with a structural fix — the parent's transformers>=5.3 and IndexTTS's transformers<5 now live in separate OS processes and can never collide. * New: backend/engines/indextts/ — sidecar package (__init__.py hosts IndexTTS2Backend, main.py is the sidecar entrypoint, bootstrap.py owns the 3-step venv probe + lazy uv-based bootstrap). * services.tts_backend: IndexTTS2Backend's in-process body removed; registry resolves the class lazily via a _LazyRegistry indirection + PEP 562 __getattr__ re-export. This breaks the import cycle that arose when both subprocess_backend and tts_backend tried to import each other at module load. * docs/engines/indextts.md: install walkthrough + venv resolution order + common errors (linked from is_available()'s unavailable message). * tests: - test_indextts_backward_compat.py (8) — probe priority, no-spawn discipline, HF cache marker preservation (ENGINE-07). - test_indextts_sidecar.py (17) — subclass shape, isolation_mode, parent-side emotion arbitration (vector/audio/text/description), coexist-with-OmniVoice (headline #42 closure), env forwarding. - tests/fixtures/mock_indextts_sidecar.py — stdlib-only sidecar mimicking the production wire protocol; emits 1 s sine wave. - test_issue_fixes.py: two obsolete in-process-conflict tests rewritten to assert the new subprocess contract (no indextts.* import in the parent). Hard constraints honored: backend/services/sonitranslate.py and gpu_sandbox.py are untouched (D1 / D4). Existing v0.2.7 users with OMNIVOICE_INDEXTTS_DIR and a populated HF cache reach a working generation with zero re-download and zero re-install. 44 tests pass across the four exercised files. Full suite: 391 passed, 10 skipped, 13 xfailed, 1 xpassed in 57 s. Smoke: 4 passed. Closes #42. Requirements: ENGINE-02, ENGINE-03, ENGINE-04, ENGINE-07. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
766e2f7284 |
Phase 0 — Gates: cross-platform CI matrix + regression fixture + release smoke (#71)
* docs: initialize OmniVoice stabilization milestone project * chore: add project config (yolo + balanced) * docs: domain research for stabilization milestone * docs: define v1 requirements for stabilization milestone * docs: add GGUF + singing engine spike requirements (Phase 4 new) * docs: roadmap revision + CLAUDE.md (7 phases, 62 reqs, +GGUF/SING spikes) * docs(phase-0): add Gates phase RESEARCH.md Phase 0 research synthesizes the cross-platform CI matrix, frozen omnivoice_data fixture, installer post-build smoke, SHA-256 checksum publishing, and PR-template extension into copy-paste-ready YAML and Python snippets composed entirely from existing in-repo patterns. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(phase-0): add Gates phase CONTEXT, PATTERNS, and PLAN Phase 0 — Gates is the hard pre-condition for v0.3.x stabilization. Lays cross-platform CI matrix (macos-14/windows-2022/ubuntu-22.04), regression fixture (≤200 KB), installer smoke on tag push, SHA-256 checksums in release body + per-OS SHA256SUMS-*.txt assets, PR template with RC cadence + fixture line, and the open-PR landing for #51. Plan covers GATE-01..06; structured into 7 slices (A–G) with explicit Slice C → Slice G dependency reordering so the new smoke-matrix lands on main before PR #51 (CONTEXT.md L86 interleave decision). Plan-checker iteration 2: APPROVED — all 3 BLOCKERs + 3 MAJORs from iteration 1 resolved (file truncation/Slice-G missing, GATE-06 sibling PR verification, Slice C ordering, Truth #5 wording, macOS Tauri WebView avoidance per Pitfall #5, Windows taskkill per Pitfall #2). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(00-gates): seed regression fixture (GATE-01) - scripts/seed-test-fixture.py — deterministic builder for tests/fixtures/omnivoice_data/ - wipes + rebuilds; fixed created_at=1700000000.0; all-zero PCM for byte-deterministic diffs - calls backend.core.db.init_db() directly (alembic versions/ is empty — see CONTEXT.md) - checkpoints WAL → DELETE on close so no -shm/-wal sidecars pollute git status - exits non-zero if fixture > 200 KB - tests/fixtures/omnivoice_data/{omnivoice.db, README.md} — 8-table empty DB + 1 voice_profiles row - tests/fixtures/omnivoice_data/voices/test-voice/{profile.json, sample.wav} — 1-sec 24 kHz mono silence - .gitignore — explicit allow-list (!tests/fixtures/omnivoice_data/**) so the existing omnivoice_data/, *.db, *.wav patterns don't hide the fixture from git Verifies: du = 144 KB on disk; sqlite_master lists 8 init_db tables + sqlite_sequence; voice_profiles has exactly 1 row id='test-voice'; 0 rows in generation_history. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(00-gates): add tests/smoke/test_boot_smoke.py (GATE-01) - tests/smoke/__init__.py — package marker so pytest treats tests/smoke/ as a module - tests/smoke/test_boot_smoke.py — 4 in-process FastAPI TestClient smoke tests: * test_health_returns_ok — /health returns 200 + {status:ok, device:...} * test_profiles_endpoint_lists_fixture_voice — /profiles surfaces the seeded test-voice row (validates OMNIVOICE_DATA_DIR wiring → DB_PATH → init_db schema) * test_system_info_includes_data_dir — /system/info resolves data_dir * test_history_endpoint_empty — /history reaches DB and returns [] Test isolation env vars (OMNIVOICE_MODEL=test, OMNIVOICE_DISABLE_FILE_LOG=1) set at module top BEFORE any backend import — pattern from tests/test_router_smoke.py. Fixture is copied to a per-session temp dir so the test never mutates the checked-in artifact (SQLite file-change counter + runtime subdirs like dub_jobs/ would otherwise dirty `git status` after every run). Failure mode: if tests/fixtures/omnivoice_data/ is missing, pytest.fail at import time with the regenerate command. - .gitignore — tighten the GATE-01 allow-list to ONLY the seed-produced files (README.md, omnivoice.db, voices/test-voice/profile.json, sample.wav). Prevents future runtime subdirs the backend may create under the fixture from being accidentally committed. Verifies: `uv run pytest tests/smoke/ -q --tb=short` → 4 passed in 1.31 s (target was < 30 s). `git status` clean after a test run. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(triage): record post-planning GitHub state — PR #62, new issues, OOS deferrals - GATE-06: mark #53 + #61 merged (2026-05-16); add #62 (Wave 1 quick wins) to gate set - INST-01: note PR #62 implements setuptools pin (closes #58) - INST-04: note PR #62 lands README docs for #56 workaround - INST-12: new requirement for #65 Windows Triton/torch.compile OOM (filed post-planning) - Out of Scope: defer #67/PR #68 (audio effects), #64 (custom model dir), PR #66 zh-CN (i18n milestone), #63 (empty-template bug) PR #62 is the user's own Wave 1 work landed as a separate PR while GSD planning ran in parallel. Merging it eliminates duplicate work in Phase 1. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci(00-gates): add cross-platform smoke matrix (GATE-02) - New smoke-matrix job on macos-14, windows-2022, ubuntu-22.04 - needs: test, fail-fast: false, timeout-minutes: 10 - Pinned actions: checkout@v4, setup-python@v5, setup-uv@v3 (cache enabled) - Per-OS ffmpeg + libsndfile install (brew/choco/apt via awalsh128 cache) - UV_HTTP_TIMEOUT=120, UV_HTTP_RETRIES=5 for restricted-network resilience - Narrow scope: uv run pytest tests/smoke/ -q --tb=short - Existing `test` and `tauri-cross-platform` jobs untouched Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci: add workflow_dispatch to ci.yml so smoke-matrix can run on feature branches * feat(00-gates): add --health-check CLI flag to backend entrypoint (GATE-03) - argparse on __main__ block; --health-check boots uvicorn in a daemon thread and polls http://127.0.0.1:3900/health every 5s for up to 60s. - Prints 'OK — /health responded 200 after Ns' and exits 0 on first 200. - Prints 'FAIL — /health did not respond 200 within 60s' to stderr and exits 1 on timeout. Default invocation behavior unchanged. - No new deps (stdlib argparse/threading/time/urllib.request/sys + uvicorn). - Consumed by per-OS installer-smoke step in .github/workflows/release.yml. Verified locally: exits 0 in 5s against tests/fixtures/omnivoice_data/. * ci(00-gates): add per-OS installer smoke to release.yml (GATE-03) Adds three matrix-leg-specific steps after 'Build + release (Tauri)', each gated by runner.os with timeout-minutes: 5: - macOS (macos-14): hdiutil attach DMG → locate bundled Python backend inside *.app/Contents (NOT the Tauri WebView shell — RESEARCH Pitfall #5: WebView hangs on headless runners) → invoke --health-check → hdiutil detach. Falls back to *.app/Contents/Resources and hard-fails with a directory listing if no backend binary found. - Windows (windows-2022): msiexec /quiet install → find backend.exe under 'C:/Program Files/OmniVoice Studio' → invoke --health-check in background, wait, then taskkill //F //T //PID to cleanup orphaned PyInstaller child processes on port 3900 (RESEARCH Pitfall #2). - Linux (ubuntu-22.04): --appimage-extract (no FUSE on GH runners), locate binary or AppRun, run under xvfb-run -a. Bundle-only regressions (PyInstaller missing-module, Tauri sidecar path mismatch) are invisible to ci.yml's in-process smoke matrix — this step closes that gap before any release is published. Verified: YAML parses; all three steps present; gating + timeout correct; Pitfall #2/#5 mitigations preserved. * ci(00-gates): publish SHA-256 checksums in release body + as asset (GATE-05) - Add 'Compute SHA-256 checksums' step writing SHA256SUMS-<label>.txt per matrix leg using native shasum/sha256sum (Git Bash on Windows). - Add 'Append checksums to release + attach SHA256SUMS file' step using softprops/action-gh-release@v2 with append_body: true so the hashes land in the release body alongside tauri-action's content (not replacing it) and the file is uploaded as a release asset for 'shasum -c SHA256SUMS-<label>.txt' verification. - Both steps gated by 'github.event_name == push && refs/tags/v*' so workflow_dispatch dry-runs do not attempt to attach to a non-existent release (per CONTEXT.md L70 + RESEARCH Pitfall #7 deferral of any aggregate cross-leg SHA256SUMS job). - fail_on_unmatched_files: true to surface path-resolution errors loudly. * docs(00-gates): document RC cadence + regression-fixture check in PR template (GATE-04) * docs(setup): add HF token persistence guide for macOS/Windows/Linux (DOCS-05) Covers two persistent paths: - Method A — canonical ~/.cache/huggingface/token via huggingface-cli login - Method B — shell env var (~/.zshrc / ~/.bashrc / Windows User scope) Documents the v0.2.7 "session only" in-app behavior + notes that Phase 1 AUTH-03 will make in-app pastes write to the canonical file. Bundled with Phase 0 PR per user request. Strictly DOCS-05 scope — zero code changes, no engine touches. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * spec(auth): redesign HF token resolution as 3-source cascade with fallback (AUTH-01..06) Replaces the env_store.py file-based design with a SQLite-backed app store + cascade resolver that checks app → env var → ~/.cache/huggingface/token in priority order, with automatic fallback to next source on HTTP 401. User-explicit design decision: - App-stored token (SQLite settings table, AES-GCM encrypted) wins - Env var ($HF_TOKEN) second - Global huggingface-cli login file third - All three sources visible in Settings → API Keys with "Active" badge - Save action populates BOTH app store AND canonical HF file (defense in depth) New requirement: - AUTH-06 — on 401, auto-retry next source in cascade before erroring Also: traceability count corrected (62 → 74 — undercount at planning + INST-12 + AUTH-06 added post-planning). All 74 v1 reqs mapped. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(auth): backend recognizes HF token from canonical file, not just env var Two call sites were only checking $HF_TOKEN env var, missing the canonical ~/.cache/huggingface/token file written by `huggingface-cli login` (or the app's future Save action): - system.py `/system/info` `has_hf_token` flag — UI showed "No HF token" even when `huggingface-cli login` had populated the file. - model_manager.get_diarization_pipeline — pyannote diarization silently returned None when only the canonical file was set. This is the bug behind issue #35 (speaker diarization setup failure). Both fixes use the same pattern: env var > huggingface_hub.get_token() (which reads the canonical file). Adds a local _has_hf_token() helper to system.py with a comment marking it as prelude to the AUTH-01..06 cascade (Phase 1 token_resolver.py will layer SQLite app-store on top). Closes #35 sub-issue (canonical token invisible to diarization). Cross-cuts AUTH-02 + AUTH-06 design for Phase 1. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(dictation): make pill-widget mode reachable from GUI + scripts (INST-13) The dictation widget infrastructure shipped in PR #40 but was only reachable via the undocumented --pill CLI flag. Adds three discovery paths: 1. Tray menu: "Switch to Dictation Widget" (studio mode) — saves launch_as_widget=true to config, relaunches with --pill, exits current. Mirrors the existing "Open Studio" path in pill-mode tray. 2. Persistent config: AppConfig.launch_as_widget (bool, default false). Read at startup via load_config_pre_app() (uses dirs-next, no AppHandle required). CLI --pill still takes precedence when explicitly passed. 3. Tauri commands: get_launch_as_widget / set_launch_as_widget for the Phase 2 Settings UI to bind a checkbox to. 4. Scripts: bun desktop-prod:pill / desktop-prod:run:pill — forward --pill to the bundled app launch. macOS uses `open -n --args` to spawn fresh instance with the flag. Closes the GUI half of INST-13. Phase 2 closes the Settings UI half. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(dictation): show widget unconditionally on pill-mode launch + visible Suspense fallback Before: pill mode set up correctly but the widget window stayed hidden until ⌘⇧Space was pressed. New users saw absolutely nothing on launch (no main window, no dock icon, hidden widget) and assumed the app failed. If global-shortcut Accessibility permission wasn't granted, they had no path to discover the widget at all. Two changes: 1. lib.rs: in pill_mode_setup, explicitly show + position + focus the widget window after hiding main. With per-call error logging so we can diagnose failures (and a clear error log if widget window wasn't created at all — points at tauri.conf.json regression). 2. main-app.jsx: Suspense fallback was `null`, which combined with widget's transparent+decorations:false config made any lazy-import delay or failure invisible. Now renders a dark pill saying "Loading dictation…" so even if CaptureWidget lazy-import stalls, the user sees the window exists. Studio mode behavior unchanged — widget stays hidden until hotkey or tray click triggers it (existing show() call in the shortcut/ menu handlers is preserved). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(dictation): create widget window programmatically; Tauri 2 silently dropped config-array creation Root cause: declaring the widget window in tauri.conf.json's app.windows[] silently failed in Tauri 2 — get_webview_window("widget") returned None even though the config was syntactically valid. Probable culprit was the transparent + decorations:false + visible:false combo, but Tauri offered no error message either at startup or via webview_windows() enumeration. Diagnosed by adding webview_windows() enumeration logging at setup start (only ["main"] ever appeared) and a programmatic WebviewWindowBuilder fallback that surfaces real Result errors. Fix: - tauri.conf.json: widget entry now has `create: false` to make the config-vs-programmatic handoff explicit. - lib.rs setup(): call WebviewWindowBuilder::new(app, "widget", ...).build() with the exact same surface attributes the config used to declare. - capabilities/default.json: include "widget" in windows array so the new window inherits the same Tauri permissions as main. - tauri.conf.json: remove the invalid `"url": "/?window=widget"` field — WebviewUrl::App takes a path only, query strings aren't supported. Both windows now load index.html. - main-app.jsx: replace URL-query-based widget detection with getCurrentWindow().label === 'widget' via @tauri-apps/api/window. This is the Tauri 2-recommended pattern for multi-window apps and works regardless of URL routing. Closes the immediate UX bug behind the dictation widget being invisible. Builds cleanly + manually verified: pill widget visible on screen at top-center after `bun desktop-prod:pill`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
67328d04fe |
refactor: split backend into api/core/services/schemas, harden security + fd pressure, add searchable language picker, fix segment fragmentation
Backend:
- Split monolithic main.py into backend/{api/routers,core,schemas,services}
- core/db.py: allowlist-gated migrations, db_conn context manager (kills SQL injection on ALTER)
- core/tasks.py: lock-guarded listener add/remove/push, snapshot-before-iterate
- services/ffmpeg_utils.py: run_ffmpeg helper with concurrency semaphore, EAGAIN retry, guaranteed reap
- services/segmentation.py: Bengali/CJK/Arabic punctuation, ultra-short tier, stitch_adjacent_shorts,
bounded-loop merge; public clean_up_segments API
- services/model_manager.py: robust lock.locked() handling
- api/routers/dub_core.py: job_id traversal guard, thread-safe _active_procs, timeouts on ffmpeg/demucs,
POST /dub/cleanup-segments endpoint
- api/routers/dub_export.py: guarded SSE listener remove, ffmpeg timeouts via run_ffmpeg
- api/routers/exports.py: destination_path validation, safe source resolver, subprocess list-form
- api/routers/generation.py: contextlib.suppress on tempfile cleanup, db_conn usage, safe output-path helper
- api/routers/system.py: try/finally tmp cleanup, subprocess timeouts
- schemas/requests.py: TranslateSegment.id int->str to match hex segment IDs
- main.py: threading.Lock around crash log writes
Frontend:
- components/SearchableSelect.jsx: popover combobox with search, keyboard nav, popular+recent pins, 200-item cap
- App.jsx: wire SearchableSelect for dub language / ISO code / voice-gen language; Clean Up segments button;
fix blob URL leak (object-shaped prev in setter, unmount cleanup via ref)
- components/WaveformTimeline.jsx: explicit <video> detach instead of innerHTML='' to release decoder
- index.css: ss-* combobox styles matching Gruvbox theme
Tests:
- tests/test_segmentation.py (26 cases), test_dub_transcribe.py, test_dub_export_unique.py, conftest.py
Chore:
- .gitignore: exclude omnivoice.zip, /research/ reference clones
- Remove tracked stray root test scripts + crash_log.txt
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|