Commit Graph
118 Commits
Author SHA1 Message Date
psiberfunk fb98bbfaf6 fix(audiobook): surface interrupted renders 2026-09-07 21:55:40 -04:00
Palash Debnath 157f2987dc fix(api): retain Node ESM compatibility for abortable delay imports 2026-09-07 11:17:24 +05:30
Palash Debnath 6c47ae7b21 fix(api): honor cancellation throughout transport diagnostic waits 2026-09-07 11:03:05 +05:30
Palash DebnathandClaude Opus 5 31d90def83 fix(errors): don't claim the backend crashed with no evidence that it did
Two Apple Silicon reporters were told "it most likely crashed or was killed
mid-request" while generating. Neither bug report carried a crash marker,
because none had been recorded — the app had no evidence for the one thing
it asserted, and the advice that follows that sentence is Retry and Clean &
Retry, which rebuilds the whole Python environment to fix a backend that had
not died.

Two causes, both fixed here.

The desktop shell learns the backend died from a ~2 s poll: it has to notice
the child exit before it can write the marker. `apiFetch` asked for that
marker exactly once, at the instant the transport gave up, so it raced the
poll and lost either way round — a backend that really died was reported
with the vague sentence instead of its exit code and crash notice, and one
that never died was reported as dead anyway. `streamDropError` already waits
that poll out (#1119); the request path never did. The loop is now a shared
`awaitBackendCrashMarker`, used by both, with a shorter budget here because
the transport cascade has already cost the user a few seconds.

And the copy itself overshot what it could know. By construction it is
reached only once a crash has been looked for and not found, so it no longer
names one: it says the backend stopped answering with no crash recorded, and
that a heavy job holding the engine is the likelier story — which on a
memory-pressured Mac mid-generation it is. Updated in all 21 locales, since
a translation still asserting a crash would be the same bug in another
language.

The #1337 test that required the crash wording is updated with it: #1337
established that the backend had answered seconds earlier, not what silenced
it, and requiring the stronger claim is what pinned this in place.

Fixes #1802.
Fixes #1805.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HR6J9zKQop9TGGVUwjypnF
2026-09-04 18:03:23 +05:30
Matt Van Horn 1515b46adb feat(batch): watch-folder auto-ingest (#1768)
Opt-in watch folder on the batch queue: pick a directory once and new videos are auto-enqueued with the last Add-to-queue settings, with pause/stop controls and copy-in-progress protection. Files stream to the loopback backend as bytes; paths never leave the app. Also gives the batch queue a reachable UI entry point and streams multipart uploads to disk. Maintainer fix: the watched directory handle is opened with full share mode on Windows so users can rename or delete the folder while it is watched, matching macOS/Linux behaviour, with a cross-platform regression test. Thanks @mvanhorn!
2026-09-03 18:47:32 +05:30
Matt Van Horn 999345de41 feat(dub): realtime dub preview (#1769)
Opt-in live preview for dub segments: edits debounce into a streamed /ws/tts synthesis played through the chunk player, with cancellation preserved through buffered playback. Maintainer fixes: /ws/tts added to the backend ticket allowlist (feature was dead off-loopback), handshake failures surface a toast, loopback-only plaintext refusal reverted to keep the documented remote-GPU setup working, PCM16 decode hardened. Thanks @mvanhorn!
2026-09-03 18:27:07 +05:30
Matt Van Horn a95041f1e6 feat(studio): speech-to-speech voice changer (#1765)
Add a bounded, local-first speech-to-speech Convert workflow with shared ASR/TTS admission, duration matching, stale-request cancellation, profile conditioning, watermarking, persistence, localization, and regression coverage.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
2026-09-02 09:45:40 +05:30
Matt Van Horn 4053397921 feat(dub): karaoke word-highlight caption burn-in (#1764)
Adds opt-in word-timed ASS karaoke captions while preserving the existing line-caption default.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
2026-09-02 08:32:43 +05:30
Palash Debnath 497d57ee62 Show complete engine disk costs before install (#1728)
Closes #1718.

Adds structured pre-install and measured post-install disk costs, complete sidecar preflight accounting, strict authorization for recursive scans, and localized catalogue details with confidence and deduplication context.
2026-09-01 23:49:05 +05:30
Palash Debnath a4f68e4ea0 Preserve cloned voices during queued SRT imports (#1745)
Closes #1709.

Queues SRT selection until speaker analysis and clone extraction finish, then applies the newest selected subtitle file with stale-result, failure, replacement, retry, and abort guards.
2026-09-01 23:15:43 +05:30
Palash Debnath 1103898d7b Merge remote-tracking branch 'origin/main' into fix/runtime-stability-1652
# Conflicts:
#	CHANGELOG.md
2026-08-27 22:30:58 +05:30
Palash Debnath 57263edad7 fix(tts): centralize single-generation admission 2026-08-27 20:38:29 +05:30
Palash Debnath 1d445855d2 fix(dub): make language workflow reliable 2026-08-27 19:43:49 +05:30
Paolo Antinori 871d68a6ff fix(auth): offer API-key login on server-mode admin 403s (#1569)
Fix server-mode admin authentication recovery without trapping PIN-only deployments, and prevent stale 403 responses from clearing or superseding newly issued sessions. Includes backend/frontend regression coverage, docs, and changelog credit for @paoloantinori.
2026-08-20 00:03:37 +00:00
Palash DebnathandClaude Fable 5 0ee62b2261 feat(gallery): save gallery voices as profiles, with validated audio references (#1542)
* feat(gallery): save gallery voices as profiles, with validated audio references

Work-in-progress lifted from the concurrent gallery session at the
owner's request (its uncommitted working tree, preserved verbatim from
base 92b1ee5d; safety snapshot remains at rescue/gallery-wip):

- gallery voices can be saved as local profiles: audio is copied into
  the profile store with content-addressed filenames, existing profiles
  are detected and refreshed only when the source clip changed
- backend/core/audio_validation.py: symlink-rejecting, root-contained
  resolution for persisted profile WAV references, with tests
- archetype/community routers and the Voice Gallery UI updated for the
  save-as-profile handoff (spec: docs/specs/longform/26-gallery-use-handoff.md)
- locale updates for the new gallery strings across all 21 files

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: drop a stray local screenshot script that rode in with the tree copy

* fix(community): explain the tolerated Content-Length parse failure; drop an unused import

CodeQL on #1542: the empty except now says why it is safe (the streamed
byte counter enforces the same cap regardless), and the test file loses
an unused Path import.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(gallery): review findings — copy outside the write lock, no stale completions

CodeRabbit on #1542, all findings addressed:
- the profile-audio copy stages to a .part temp BEFORE BEGIN IMMEDIATE
  and publishes via atomic os.replace inside it — other backend writers
  no longer block for the duration of an audio copy; a mid-copy failure
  leaves no temp droppings and no profile row (both pinned by tests)
- VoiceGallery async ops carry per-operation generation tokens: a
  preview or save-as-profile that resolves after unmount (or after a
  newer operation) can no longer play audio, redirect into a workspace,
  or touch state — three fail-before regression tests
- VoiceGalleryActions imports the page at test runtime; the e2e locator
  uses a stable data-testid instead of a translated string; symlink
  tests skip cleanly where the OS can't create symlinks; the changelog
  line carries its PR ref

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ci: static ffmpeg fallback when the chocolatey feed is down

Third feed outage to break a PR run (2026-07-20, 2026-07-28, today —
three attempts, three 'installed 0/1'). Chocolatey is a distribution
channel, not the dependency: after the retry loop exhausts, fetch the
static gyan.dev build from its GitHub release mirror and put it on
PATH — same binary, no feed in the path. URL verified live (HTTP 200).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 22:48:05 +00:00
19ae20111a fix(security): replace persistent admin keys with scoped sessions (#1528)
* fix(security): replace persistent admin keys with sessions

Exchange the remote administrator key once for bounded, revocable credentials. Canonicalize backend principals, enforce cookie CSRF and exact origins, and use path-bound one-use WebSocket tickets.

Migrate the bundled UI away from durable master-key storage and credential-bearing URLs. Add unit, integration, static-hygiene, and production-browser regressions plus synchronized operator documentation.

* docs: link session hardening to PR 1528

* fix(security): key session indexes with process pepper

Use HMAC-SHA-256 instead of an unkeyed digest for in-memory session and WebSocket-ticket indexes. This preserves constant-size lookup identifiers, makes copied records unusable without the process pepper, and resolves CodeQL's weak sensitive-data hash finding.

* fix(auth): align empty bearer migration precedence

Centralize the Authorization-channel presence decision with canonical principal parsing. Bearer followed only by spaces now remains an empty channel during legacy cookie migration, while unsupported or invalid explicit credentials stay authoritative and fail closed.

* fix(security): harden admin session review boundaries

* fix(security): derive key generations with HKDF

* fix(auth): anchor the admin-session store so module reloads cannot fork it

test_master_exchange_does_not_bypass_pin_on_normal_routes failed in full-suite
runs: test_mcp_bindings' client fixture purges the services.* tree from
sys.modules and reloads main, so api.routers.auth re-imported a fresh
services.admin_sessions (new AdminSessionStore) while core.auth kept its
import-time reference to the old one — the exchange issued the cookie into
one store and the middleware resolved it against another, turning the
expected "PIN required" into "API key required".

Root cause is the class of bug, not the one test: a process-global auth
store defined as a bare module-level singleton forks under importlib.reload
or purge-and-reimport. Fix at the source: admin_session_store now resolves
through a synthetic sys.modules anchor (_omnivoice_admin_session_store_anchor)
that reloads never re-execute and package-prefix purges never match, so every
copy of the module shares the one per-process store. No consumer or behavior
changes.

Regression test reproduces both fork vectors (in-place reload and
sys.modules purge + fresh import) and asserts previously issued sessions
still resolve and the store identity is preserved; it fails before this fix
and passes after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(auth): honor X-Forwarded-Proto for CSRF origin and Secure cookies behind TLS proxies

Behind Tailscale Serve (docs/remote-gpu.md) or any TLS-terminating proxy,
the browser talks https while the backend hop stays http, so exact-origin
CSRF compared an https Origin against an http expectation and rejected
every legitimate request, and the session cookie shipped without Secure.
uvicorn's ProxyHeadersMiddleware only rewrites the scope for loopback
peers, which misses Docker and any non-loopback proxy topology.

New core.csrf.effective_scheme derives the client-facing scheme: resolved
scope first (uvicorn's trusted-proxy rewrite wins), then an upgrade-only
read of X-Forwarded-Proto's first value — https/wss promotes http to
https, everything else is ignored, and a genuine TLS hop can never be
downgraded. Used by both the destination-origin comparison and
auth._secure_cookie so the WS-ticket/logout CSRF paths and the cookie
Secure flag agree. Spoofing gains nothing: the host:port half of the
origin tuple is untouched, browsers cannot attach the header cross-site
without a preflight this API never grants, and forging it on plain http
only adds Secure (the browser then drops the cookie — self-harm only).

Regression tests: proxied https origin accepted (origin check, Secure
flag, logout), comma-separated chains, scope-fallback path, spoofed
header still rejects cross-origin, cannot downgrade real https, junk
values ignored.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(auth): consume the stored admin key only after a successful exchange

A remote-backend user upgrading with their backend unreachable lost the
only stored copy of OMNIVOICE_API_KEY: every migration path deleted the
durable ov_api_key BEFORE the session exchange settled, stranding them
until they recovered the key from the server box. Close the whole class:

- client.ts bootstrap: read the legacy key, exchange first, and remove
  the durable copy only after the exchange succeeds; on failure the key
  stays so the next launch retries the migration (auth gate still rises).
- authSession.ts exchangeApiKey: move removeLegacyMaster from before the
  fetch to the cookie/bearer success paths — the key never coexists with
  a live session, but a rejected or hung exchange no longer consumes it.
- remoteBackendProbe.ts configuredRemoteBackend: stop wiping the key on
  every app mount.
- RemoteBackendPanel: a connection test or an aborted save no longer
  wipes the pending key; only disabling the remote backend discards it.
- prefKeys.js: ov_api_key moves from PREF_KEYS to PRESERVED_KEYS —
  factory reset preserves the pending connection credential exactly like
  ov_backend_url; the successful migration is what deletes it.

Tighten the credential-hygiene static guard to match: it accepted
sessionStorage.setItem('ov_api_key', …) — the exact class it exists to
close. The guard now flags .setItem(<master key>) on any storage
receiver, quote style, or injected-store alias, with a self-test pinning
what it catches and what stays legal.

Fail-before/pass-after regression tests: backend unreachable retains the
key and the next bootstrap retries it; a successful exchange removes it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* perf(auth): make session validation occupancy-independent

* test(auth): catch optional master-key storage calls

* feat(docs): add PR control document for bultodepapas in VoiceStudio

* docs: keep the PR tracking board in the fork; credit the changelog line

The pr-control document is excellent process discipline, but it is the
contributor's own operational board (their inventory, their update
commands) — it lives naturally in their fork, and docs/agents/ here is
context every repo agent loads. Removed with appreciation; the changelog
line gains its contributor credit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 20:46:24 +00:00
Palash DebnathandClaude Fable 5 9832fbd693 feat(engines): switch engines from anywhere — footer quick switch, workspace chips, shortcuts (#1530)
* feat(engines): add quick switching controls

* docs(changelog): note engine quick switching (#1530)

* fix(support): theme amount cards

* fix(support): restore themed amount cards

* fix(engines): address quick switch review findings

* fix(engines): green the full frontend suite around the quick switch

Three failure classes the targeted runs missed:
- the popover referenced --chrome-radius, which does not exist; it now
  wears the footer's shared MENU_SURFACE like the compute popover
- LogsFooter tests hand-wrote their api/system and api/hooks mocks, which
  drop every export the footer gains next; they are partial mocks now
- DubHeader/AudiobookHero tests rendered without a QueryClientProvider,
  which useEngines needs

Also: workspace-header chips open the popover downward (dropUp stays on
the footer instance) so it cannot clip off the top of the viewport.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(engines): one QueryClient per test module, not per render

CodeRabbit: the inline client made every wrapper render a fresh cache,
so rerender() restarted the /engines query mid-test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-13 15:21:05 +00:00
Palash DebnathandClaude Opus 5 41722afe3b refactor(launchpad): quieter, borderless design refresh (#1515)
* refactor(launchpad): quieter, borderless design refresh

The launchpad carried decoration from an earlier direction: icon chips,
corner-hung count badges, a permanently visible filled arrow, uppercase
mono card titles, and a dotted stipple divider — plus a frame that had
been invisible since the app-wide border tokens were zeroed.

Rework it around what the borderless direction actually implies:

- Feature tiles get a whisper-faint surface instead of a dead frame, and
  read as three bands (bare glyph + count / title + arrow / description).
  `--card-hue` is spent sparingly — the glyph at rest, the surface, count
  and arrow only once raised. Titles move to sans sentence case; counts
  are plain tabular numerals. Lift softened 4px -> 2px, coloured glow ->
  neutral shadow, plus an explicit focus ring and a staggered entrance.
- Hero drops the boxed "646" pill and the filled A/B-Compare button for
  quiet type, with a hairline standing in for the separation.
- Section labels trade the dotted stipple for a single fading hairline;
  rows are transparent until hover and reveal "Open" on hover/focus (it
  stays in the DOM, so AT and keyboard always reach it).
- Hero, tiles, recent files, callout and project lists now share one
  1180px column — previously only the top half was capped, so lists ran
  edge-to-edge on a wide display while the deck stayed centred.

Two bugs found and fixed while doing it:

- Buttons that had `border border-solid border-transparent` removed fell
  back to the UA default border and rendered a visible 1px outline. They
  now carry `border-0` explicitly.
- `.lp-animate` used `animation-fill-mode: both`, so after the entrance
  it kept owning `transform` — and animation-origin declarations outrank
  normal ones, which silently killed the card hover lift. Now `backwards`,
  which still holds the from-state through the stagger delay.

Also drops CSS the page has not rendered since #904: the cursor-spotlight
layer, the breath ring, and the per-card waveform strip.

Verified with headless renders at 1600/1280/940 and the empty state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(dictation): decode Wayland portal signals and show the capture pill

The GlobalShortcuts portal declares Activated/Deactivated as
(o session, s shortcut_id, t timestamp, a{sv} options). We decoded the
timestamp as u32, so zbus rejected every signal with

  Signature mismatch: got `(osta{sv})`, expected `(osua{sv})`

and the press was dropped as an invalid signal. Registration succeeded
and the desktop even reported the bound chord back, so the hotkey looked
wired up while doing nothing at all — on every Wayland compositor, for
the whole life of the feature (#1490). Decode the 64-bit timestamp, and
keep the 32-bit spelling as a fallback so a non-conforming portal
degrades to working rather than to silence.

With presses arriving, the second half of the failure showed: nothing
had shown the widget window since it became a hidden recorder host, so a
capture ran with no pill on screen — and a mic or Accessibility failure
rendered into a window nobody could see. Add show_dictation_pill, which
bottom-centres the capsule on the monitor under the pointer and shows it
without taking focus (Windows keeps SW_SHOWNOACTIVATE so paste still
lands in the user's document), and call it from the widget for every
state but idle. Wayland denies clients their own placement, so the
compositor picks the spot there; the pill still appears.

dispatch_dictation_capture now logs whether a press was emitted or
queued — a press that reaches Rust and produces nothing was otherwise
indistinguishable from one the compositor never delivered.

Tests: portal signals decode at both timestamp widths (the 64-bit case
fails before this change with the exact production error); pill
placement centres, respects a second monitor's origin, and clamps rather
than going off-screen; the widget shows for a state needing the user,
stays hidden while idle, and never shows for a press that arrives while
dictation is disabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: sync in-progress workspace changes

Uncommitted work already in the tree, checkpointed so the branch matches
the local machine:

- Remote GPU workers: join-from-the-app flow, one-time secrets, QR join
  codes, a Compute control in the status bar, and the device-list
  Workers panel (#1516)
- Model Catalogue workspace, with Settings pointing at it
- Settings sidebar search and keyboard navigation
- Demo assets for dubbing, dictation and voice design, plus the scripts
  that render them
- Backend: validation-error handling, ASR request-path degradation, and
  the accompanying tests
- CHANGELOG entries for the above and for the Wayland dictation fix

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(tests): follow Engines to the Model Catalogue, and green the sweep

- test_supertonic3 asserted the license gate points at "Settings" while
  the engine now names Model Catalogue → Engines, which is where the
  accept button actually lives. The assertion follows the move; what it
  pins is unchanged — the hint must name a place the user can reach it.
- Carries the CJK allowlist entries for the rendered dub bundle (#1517)
  and the regenerated route snapshot for /workers/agent (#1516), both of
  which this branch inherits from the workspace sync.
- docs/install/linux.md: the dictation capsule is bottom-anchored
  everywhere except Wayland, where the protocol gives applications no
  say in their placement. Documented rather than left as a surprise
  (CodeRabbit).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: stop a flaky dependency fetch from failing green runs

en-core-web-sm resolves to a direct GitHub release URL, and github.com
intermittently answers `http2 error: refused stream before processing
any application logic`. uv's own three retries all land within the same
few seconds and fail together, so the whole job dies on a dependency
that has nothing to do with the change under test — it cost #1518 and
#1517 an otherwise-green run tonight.

Two changes: back off between whole `uv sync` attempts, which is what
actually clears it, and pass --no-sync to the pytest steps. `uv run`
re-resolves the environment before running, so every test step was a
fresh chance to hit the same fetch even though the install step had
already synced — that is exactly how #1518 failed, in the isolated
backend/tests step, with all 5467 tests already passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: one retry seam for every uv sync, not just the job that failed last

en-core-web-sm resolves to a direct GitHub *release* URL rather than a
package index, and github.com intermittently answers `http2 error:
refused stream before processing any application logic`. uv's own
retries all land inside the same ~10 seconds and fail together, so a job
dies on a dependency unrelated to the change under test. Tonight that
cost four otherwise-green runs across #1515, #1517 and #1518 — and the
first fix only covered the Tests job, so the next failure simply moved
to Smoke (Linux), which syncs separately.

The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh
backs off between whole attempts (15s, 45s, 90s) and every workflow that
syncs now goes through it — ci.yml (tests + the platform matrix),
release.yml, security.yml, evals.yml. It still fails loudly after four
attempts, so a genuinely broken lockfile is not disguised as a flake.

The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke
matrix has always set, which is part of why it was the one that kept
dying; it has them now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ci): pin the Intel-Mac contract by intent, not by command spelling

test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted
the literal line `run: uv sync --extra pockettts`, so routing every sync
through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The
contract it exists to protect is that the pockettts extra installs ONLY
on backend_supported legs — which the regex now pins, while leaving how
the sync is invoked free to change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: keep every uv run out of the resolver, and bound the retry budget

CodeRabbit, #1517:

- `uv run` re-resolves before running, so the smoke suite, the
  worker-artifact tests, the release test run and the eval run were each
  a fresh chance to hit the flaky direct-URL fetch outside the retry
  loop. All of them pass --no-sync now; the environment is already
  synced by the step that owns the retries. security.yml's
  `uv run --with pip-audit` is deliberately left alone — it layers an
  ephemeral package rather than running the project's own tests.
- The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a
  120 s timeout on the smoke matrix). Three attempts and 60 s of total
  backoff outlast the refusals actually observed while staying well
  inside the jobs' timeout-minutes.
- The Intel-Mac contract test pinned the smoke command literally too, so
  --no-sync tripped it exactly like the sync line did. Same fix: assert
  the contract (smoke runs only on backend_supported legs), not its
  spelling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 23:46:07 +00:00
velixio b7caa494eb feat(workers): remote downloads, audiobook chapters, and one port that stays honest
Five workstreams that finish the remote-GPU line, plus the test hole that
let a broken signature reach a commit.

**Downloads go through the normal path** (Phase 5). Rather than a second
remote-only route, the existing Models install flow became target-aware,
so a model landing on a worker uses the same code, the same progress
events and the same UI as a local one. Progress rows key on
(target, repo_id) — the aggregator keyed on bare repo_id, so the same
model downloading here and on a worker at once collapsed into one row
that told the user nothing true about either.

**Audiobooks render chapter by chapter on the worker** (Phase 8), with
per-chapter local fallback and ONE aggregated notice. The failure that
shape exists to prevent: a remote GPU that sleeps at chapter 40 of 200
must not turn a working book into 160 rows of PROGRESS_LEASE_EXPIRED.
Dictation is deliberately NOT ported — it runs ASR per utterance inside a
live WebSocket loop, and paying queue admission plus a round trip there
would spend the one thing that route is for.

**Dubbing stays local, and says so** (Phase 7). The coarse worker
operation is not finished, so the picker still reports dubbing as local
rather than showing a green remote chip over work this machine is doing.
What could not wait is the in-loop OOM retry: it sniffed the error string
and flushed the *local* CUDA cache, which under remote execution is the
wrong machine's GPU entirely. That is fixed now, before the path that
would have exercised it exists.

**Two instances can no longer share the control plane.** A second
VoiceStudio silently bound the same worker port and coexisted, so remote
workers landed on whichever process won the race — a session that
registers with one instance and appears dead to the other. This produced
hours of misdiagnosis during hardware testing and would hit any user with
the app open twice. The second instance now keeps running locally and
explains the conflict instead of quietly competing.

**And the hole that allowed all this to be missable.** gpu_gateway called
Scheduler.submit(pinned_worker_id=...) one commit before that parameter
existed. Every remote generation raised TypeError; 5236 tests passed
anyway, because nothing exercised the gateway against the real scheduler.
tests/test_gpu_gateway_scheduler_contract.py now runs that path for real
and binds every gateway→dependency call signature. Verified by renaming
the parameter away and watching both tests fail with the original error.

Gallery previews also fall back to a local render when a downloaded clip
cannot be decoded, rather than yielding silence.

Backend 5274 passed, frontend 1812 passed.

Not yet verified on hardware: Phases 4, 5, 6, 7, 8. Only the TTS path and
its artifact transport have been proven on a real GPU.
2026-08-11 16:53:33 +05:30
debpalash aeda504a03 fix(security): authorize export filesystem sinks 2026-08-10 07:32:50 +00:00
debpalash 9ee30766db Merge remote-tracking branch 'origin/main' into fix/ghas-path-boundary
# Conflicts:
#	CHANGELOG.md
2026-08-10 05:52:12 +00:00
debpalash c2b9951f28 Merge remote-tracking branch 'origin/main' into fix/ghas-path-boundary
# Conflicts:
#	CHANGELOG.md
#	backend/core/path_authorization.py
#	frontend/src-tauri/src/commands.rs
#	tests/test_network_share.py
2026-08-10 04:32:23 +00:00
debpalash 3b808ed20d fix: harden cookie import lifecycle 2026-08-10 04:01:44 +00:00
debpalash 74e6d4582b fix(dub): keep cookie credentials off remote plaintext HTTP 2026-08-10 03:16:35 +00:00
debpalash b1b18dd1ba fix(dub): support explicit YouTube cookie exports (#1429) 2026-08-09 20:42:47 +00:00
Palash Debnath 5cab8e0149 feat: rename the product to VoiceStudio (previously OmniVoice-Studio)
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.

Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:

  - bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
    macOS TCC grants, managed venv, WebView localStorage, the
    single-instance lock)
  - data directories OmniVoice / .omnivoice and omnivoice.db
  - the ~150 OMNIVOICE_* environment variables
  - the X-OmniVoice-* HTTP headers (a wire protocol)
  - the published Docker image paths
  - the OmniVoice ENGINE, which is a model name and not this product

tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.

Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
2026-08-07 01:30:58 +05:30
Palash Debnath 67ab934907 fix(desktop): stop guessing 'still starting up' when we know it crashed (#1393)
Three open issues report the same message — #1337, #1351, #1378 — and two of them captured 'Last backend response: 2 s before this report' in the very same report. A backend that answered two seconds ago is not starting up. The message told those users something its own captured data contradicted, and sent them to wait and retry instead of at the crash notice and the backend log.

#1164 built exactly this honesty for dev and server deployments and excluded desktop, which is where most users are. Desktop now goes through the same mode-aware builder: it gains the last-contact story (answered-then-stopped vs never-answered) while keeping the forensics only it has — the crash notice, Settings → Logs → Backend, Retry and Clean & Retry — which a dev or server message would have replaced with a terminal and a docker log the user does not have.

The two recovery buttons are named through i18n rather than quoted in English, since they render as Réessayer / Nettoyer et réessayer and so on.
2026-08-06 19:59:52 +05:30
Palash Debnath e953af3730 fix(client): a foreign 404 page is a routing problem, not a backend error (#1386)
A rehosted UI whose API requests land on its own static host (or a reverse proxy with no API route) got that host's 404 page back, and we echoed it verbatim — the reporter saw 'NOT_FOUND bom1::...' and had no way to know their requests were reaching the wrong server.

The backend now stamps x-omnivoice-backend on every response, exposed through CORS, so the client can tell 'the backend answered 404' from 'something else answered 404' without guessing at body shape. A 404 in any other voice is reported as a routing problem, naming the URL that answered and where to fix it. The message goes through the same i18n helper as the other backend-diagnosis copy, in all 21 locales.
2026-08-06 15:13:08 +05:30
Palash DebnathandClaude Opus 5 1600f02abf feat(privacy): build the watermark toggle the app already promised (#1308)
* feat(privacy): build the watermark toggle the app already promised

errors → enterprise_faq.a_watermark has told users "Commercial licensees can
disable it in Settings → Privacy" since watermarking shipped. That control did
not exist: /watermark/status and /watermark/settings had zero callers in the
frontend, is_enabled() passes no env= to resolve() so there was no environment
escape either, and the only way off was hand-editing prefs.json. A shipped
instruction that cannot be followed is worse than no instruction.

Adds the control, wired to the endpoints that were already there. It mirrors
AnalyticsOptIn's shape but inverts its default — analytics is OFF until you
opt in, provenance marking is ON until you opt out — and hides itself when
AudioSeal is unavailable, since an inert switch over a mark that cannot be
embedded is the same lie in the other direction.

Also corrects the FAQ string in all 21 locales: the toggle is available to
everyone, not only commercial licensees, and it only affects audio generated
after the change. (Whether disabling should be licence-gated is a product
decision — the text now describes what the app actually does.)

6 tests: reflects backend state rather than assuming it, turns off, turns back
on, renders nothing when AudioSeal is missing or the backend is unreachable,
and does not optimistically flip when the update fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(privacy): translate the watermark strings, retry the status fetch, guard unmount

CodeRabbit Major — the five new strings relied on defaultValue, so every
non-English user saw English. Translated privacy.watermark_title/subtitle/
on_toast/off_toast/failed in all 21 locales.

Greptile P1 — the status fetch was one-shot, so opening Privacy while the
backend was restarting hid the control for the rest of the session. Being
findable is the control's entire purpose (the FAQ tells people it is here), so
it now retries once after 2s before giving up.

CodeRabbit — a toggle resolving after the tab closes no longer sets state or
toasts over whatever screen the user moved to.

Tests: the two "renders nothing" cases now wait for the request to SETTLE
rather than merely to start (the initial render is empty, so they would have
passed even if the control appeared afterwards), plus a case proving recovery
from a failed first fetch. 7 passing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 14:17:45 -07:00
Palash DebnathandClaude Opus 5 5afcdf787d fix(engines): warn before a long CPU synth burns the whole budget (#1302)
* fix(engines): warn before a long CPU synth burns the whole budget

#1288 closed the under-provisioned-GPU gap but left the CPU one open, and I
missed it: a CPU-only host is a BENIGN routing verdict, so routingNotice()
correctly stays silent — yet #1299 and #1260 are exactly that shape, CPU hosts
that hit the 300s budget on long text with no warning at all. "Nothing is
misconfigured" and "this will finish in time" are different claims.

Threshold is the backend's own definition of past-short: generate_timeout_for()
gives the first 1200 characters the flat budget before extending it, so
ordinary sentences on a CPU laptop stay quiet and only the shape that actually
times out is flagged. Hardware caveats still take precedence — one toast, and
it names the real reason rather than generic advice.

5 tests; engines.cpuLongText translated in all 21 locales.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(engines): don't tell CPU-tuned engines to switch to themselves

Greptile P1. The advice names OmniVoice GGUF and Supertonic-3 as the CPU-tuned
alternatives — shown to someone already running one of them, it is advice to
switch to what they are using. Those two now get the same warning without the
self-referential clause; the engine set matches the backend's own timeout
message so the two can't disagree about who is CPU-tuned.

Also documents the preflight in docs/performance.md (docs-sync rule): both
warning shapes, why the threshold is 1200 characters (it is the figure the
budget itself uses), that they are advisory and once-per-engine-per-session,
and the CPU-tuned exception.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 06:23:30 -07:00
Palash DebnathandClaude Opus 5 07263ef42e fix(engines): warn about under-provisioned hardware before the synth, not after (#1288)
* fix(engines): warn about under-provisioned hardware before the synth, not after

Six reports are the same story: #1240, #1246, #1248, #1277, #1283, #1284 —
4 GB and 6 GB cards running an engine that wants 6 GB, each one waiting out
the full 300s compute budget to be told the job "was too heavy". The routing
layer knew the whole time. The error text even names the card and the figure.

The caveat only ever surfaced on the engine-PICK toast, so it reached people
who changed engines and nobody whose engine was already selected — the
default, or one persisted from a previous session. That is most users.
/generate does return X-OmniVoice-Routing, but a response header arrives when
the job ends, five minutes too late to be a warning.

So the check moves to the chokepoint every synth path shares (api/generate.ts,
same argument as the in-flight count). Fire-and-forget: never awaited, so it
cannot add latency to the request it warns about; never throws, so an
unreachable backend costs a warning rather than a generate; once per
engine+reason per session, so it informs instead of nagging. Advisory, not
blocking — the driver can page to system RAM and short inputs fit where long
ones don't.

Extracts routingNotice() as the single frontend mirror of the backend's
routing_notice(). Two callers now need "is this verdict worth interrupting
for", and two inline copies would drift — invisibly, until someone on DirectML
or an unavailable engine gets a hardware warning for a normal pick.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(engines): route streaming synthesis through the generate chokepoint

streamGenerateSpeech POSTed /generate via apiFetch directly — a second,
parallel door. Everything attached to "the one call every synth path shares"
therefore did not apply to it: the in-flight count that stops the updater
relaunching mid-synthesis, and the new under-provisioned-hardware preflight
(Greptile P1). A chokepoint with two doors is not a chokepoint.

Also fixes two cache defects in the preflight itself:

- An engine pick left the cached /engines response describing the PREVIOUS
  engine for up to 60s, so switching engines and generating immediately warned
  about the one you just left — or stayed silent about the one you just chose.
  notifyEngineSelected() now drops the cache, and hands over the caveat it just
  displayed so the preflight does not repeat the same sentence seconds later.
- A rejected listEngines() promise stayed cached for the full TTL, silencing
  the caveat for a minute after the backend came back. It is now evicted, but
  only if it is still the current entry, so a racing newer fetch survives.

6 new tests; 2 of the 3 streaming ones fail before this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(engines): hold the in-flight claim for the whole stream, not just headers

generateSpeech releases its claim when the Response resolves — when the
HEADERS arrive — but a streaming synth generates audio for as long as the body
is read. Routing streaming through it gave it a count for the first time, then
dropped that count to zero for the entire synthesis, so the updater saw idle
and was free to relaunch mid-stream (Greptile P1).

streamGenerateSpeech now wraps the whole operation in withTtsInflight().
Nesting is harmless because the store tracks a count, not a boolean — the
inner claim just bumps it to 2 and back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-29 01:43:39 -07:00
Palash DebnathandClaude Opus 5 c3d0d6b123 feat(update): announce updates as a toast with actions, not a wall of text (#1272)
* fix(dub): purge order was hash-dependent, and the cap could evict a live marker

main went red on my own test. Two distinct defects, both mine.

1. `targets = set(job_ids)` made iteration order depend on PYTHONHASHSEED, so
   which markers a cap-forced trim discarded was luck. That is why the test
   passed locally and failed in CI — verified: the old code passes at seeds
   0/7/42 and fails at 12345. Now a de-duplicated list in caller order.

2. The size cap could evict markers the CURRENT purge had just recorded. Those
   are the newest and the likeliest to still be held by a running job, so
   dropping one is precisely the resurrection this mechanism exists to prevent.
   A 'clear history' larger than the cap forced exactly that. The cap now never
   touches the current purge, making the real bound cap + one purge — stated
   plainly rather than implied.

Tests now pin both: identical survivors across runs, and an oversized purge
keeping all of its own markers. Verified across six hash seeds; full suite green
under 12345, the seed that reddened main.

* feat(update): announce updates as a toast with actions, not a wall of text

A version's release notes ARE the whole changelog section — v0.4.1's was 42
bullets. Any surface that renders them inline becomes unusable: older builds put
them in a blocking OS dialog that filled the screen and had to be dismissed
before the app could be touched.

Removing that dialog left the opposite failure. The only remaining signal was a
6-pixel dot beside the version number in the footer, which is easy to never
notice — so users either got shouted at or told nothing.

A toast is the middle: it names the version, offers Install and restart /
What's new / Later, and leaves. The notes stay one click away in Settings →
Updates, where there is room. Keyed by version so the 6-hourly re-check
replaces rather than stacks, and it never auto-dismisses — an update the user
hasn't answered is still true. Install declines while a generation is running,
since the relaunch would lose it.

Strings added to all 21 locales, translated rather than English-filled.

Tests pin the shape that matters: the toast takes no notes prop at all, renders
under 200 characters, and cannot stack duplicates.

* fix(update): don't relaunch while work is in flight; share the busy check

Review of #1272 found the restart guard was `dubStep === 'generating'` and
nothing else. Installing an update relaunches the process, so that permitted
throwing away a dub upload, a transcription, a translation, an export or a
standalone TTS synth. The same narrow check was written twice — in the new
toast and in UpdatesPanel — so the two could also drift apart.

Replaced with a single `isAppBusy(state)` in utils/appBusy, unioning the
signals the store actually has: the dub state machine, the floating status
pill (which every long background operation already pushes to), and
`ttsGenerating` — a new transient store field mirroring useTTS's local
`isGenerating`, which lived in a hook where no global check could see it.
`dubStep === 'editing'` is deliberately not busy: it waits on the user, and
counting it would block updates for as long as a transcript stays open.

Also from review:

- A failed lazy import of the toast fell into the outer catch, whose
  setUpdateIdle() erased the update that had just been found. The
  announcement is optional; the available state is not.
- "Dismiss" was machine-translated into the employment sense — terminate an
  employee — in de/ja/ru/zh-CN/zh-TW, on close buttons and one aria-label.
  Swept every locale rather than the four lines that were flagged.
- Hoisted the test's mock state with vi.hoisted.

CodeRabbit's changelog finding is declined: Highlights bullets carry no issue
refs by design, which tests/test_changelog_style.py enforces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(shutdown): a quit mid-generate is not a 500, and not a bug report (#1276)

#1174 made a model load interrupted by shutdown benign for the background
preload, but a *request* that triggered a load took the generic unhandled-
exception path: crash log, ERROR traceback, and an error-journal entry that
feeds the bug-report pipeline. Quitting the app with a generate queued
surfaced "500 Internal Server Error: model load skipped: backend shutting
down" and offered to file a GitHub issue for a normal teardown.

Nothing failed — the process is exiting. The handler now answers 503 with
Retry-After and an actionable detail, ahead of the crash-log/journal writes.

Frontend half of the same bug: toastErrorWithReport offered "Report" for any
error. A 503 means "not now, try again" by definition, so it now shows the
backend's message without the report action — covering a still-warming
backend too, not just this shutdown path.

Fail-before/pass-after tests on both sides, including that the shutdown case
leaves no crash-log entry and no journal record.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(models): repair a half-downloaded model whichever way it reports (#1273)

transformers has two unrelated wordings for "this snapshot has no weight
shard", sharing no words:

  hub load   "<repo> does not appear to have a file named …"
  local dir  "Error no file named model.safetensors, … found in directory …"

The self-heal and the failure classifier both matched only the first. The
second is what a load of a *subfolder* inside a cached snapshot raises —
exactly where an interrupted download leaves a half-written repo — so the
reporter got neither the automatic repair (delete broken entries →
re-download → retry) nor an actionable hint, just a raw 500. Their disk had
10.6 GB free, i.e. a download that ran out of room.

The phrase list now lives once in core.failure, so the healer and the error
text cannot disagree about what an interrupted download looks like. Both
fragments of the second wording must match — "no file named" alone is
ordinary English and must not claim the class.

Also hardens the #1276 handler found by running both suites in one session:
services.model_manager can be imported under two module names, making two
distinct ModelLoadInterruptedByShutdown classes and breaking a bare
isinstance — which silently restored the 500 this fix exists to remove. Now
matched by isinstance OR class name, with a test that raises a same-named
class from a different module.

Verified against the full tests/ + backend/tests/ session: the only remaining
failures are the four pre-existing #1269 isolation leaks, unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(update): count synths in flight at the chokepoint, not in one caller

Review round 2. Both Greptile P1s were the same weakness in my first pass:
tracking the synth in useTTS meant only the Generate tab counted (voice
previews, the compare modal, the stories editor and profile previews call
generateSpeech directly and were invisible), and a boolean meant two
overlapping syntheses cleared each other — whichever settled first reported
"idle" while the other was still running, so Install and restart discarded it.

Moved to `api/generate.ts`, around the one `/generate` call all seven paths
share, as a count with a `finally` release (so an abort or a network error
frees it too). A future synth caller is covered without opting in.
Fail-before verified: the overlap test fails against the boolean version.

Also from review:

- UpdatesPanel re-reads the busy state at click time; the render-time snapshot
  only exists to disable the button, and work can start after the last render.
- The 503 now carries the allowed-origin CORS headers. Without them the
  browser reports a bare CORS failure and the actionable detail — the whole
  point of the fix — never reaches the user. Both error responses build them
  through one helper now.
- Log `request.url.path`, not the full URL: a query string can carry tokens
  and newlines.
- `update.busy` said "finish your dub first" in all 21 locales, but the guard
  now covers uploads, transcription, translation, export and synth. Rewritten
  as work-in-progress wording, translated per locale.
- ru `common.dismissStatus` was "reject status"; missed in the earlier sweep.

Declined: CodeRabbit's request for a `(#NNN)` ref on the Highlights bullet —
Highlights carry no refs by design (CLAUDE.md), and tests/test_changelog_style.py
enforces it. Also declined gating the cache re-download behind a confirmation:
the self-heal already existed and already ran for the sibling wording, this
only stops it missing half the class, and it stays behind the existing
_hf_offline() check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(update): grey out Install while work is running

CI lint caught `busy` as unused after the click-time check moved into the
handler — and that exposed a real gap: `busy` had never been wired to
anything. The comment claimed it disabled the button; it didn't. Clicking
Install during a synth just bounced a toast back.

Now it does what it said: the Install and Restart buttons are disabled while
work is in flight, with `update.busy` as the tooltip. The click-time read
stays the authority, since work can start between the last render and the
click — the disabled state is the explanation, not the safety.

Adds the first UpdatesPanel test. The case it pins hardest is the inverse of
the bug: a busy predicate stuck at true would make the app permanently
un-updatable, which is worse than what this fixes. So it asserts enabled when
idle and during dub 'editing' (which waits on the user), disabled during an
upload, a translation, and one or more synths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 14:32:32 -07:00
debpalashandClaude Opus 4.8 6cfef5c0cf fix(routing): warn about an under-provisioned GPU before the job, not after (#1226, #1222)
Two users on 4 GB cards (GTX 1650 Ti, Quadro P2000) ran the `omnivoice`
engine, waited out the full compute budget, and were told the job "was too
heavy for the available compute … most often the GPU is VRAM-starved".

The 300s-vs-372s spread between the two reports is purely text length
(`300 + (len-1200)/40`, so 372s ⇒ ~4080 chars) — one bug, not two. Nothing
about the budget is device-aware, and nothing needs to be: the real defect is
that until the moment it failed, routing showed a clean green "accelerated".
`resolve_routing` matched on GPU *family* only, so a 4 GB card and a 24 GB
card were indistinguishable, and no engine declared a VRAM requirement
anywhere in the repo.

- `TTSBackend.min_vram_gb` — advisory metadata alongside `gpu_compat`. Only
  `omnivoice` declares one (6 GB), derived from the pool's own measured
  per-job budget (`_GPU_VRAM_PER_JOB_GB = 5.0`) plus resident weights.
  Inventing floors for engines with no measured figure would put confident
  numbers in the UI that nothing backs.
- `resolve_routing` takes the floor and emits an accelerated-with-caveat
  reason when the host is below it. Reuses the existing caveat channel, so
  the Settings matrix and the synth-time routing notice surface it with no UI
  change. Advisory, never blocking: drivers page to system RAM, and short
  inputs fit where long ones don't. Kernel-risk still outranks it, and a
  failed VRAM probe (0.0) never guesses.
- `_timeout_guidance` names the actual card and its VRAM, and leads with
  "pick a lighter engine" instead of wording that reads as transient
  contention the user can flush their way out of.

Regression test: tests/test_low_vram_advisory.py (8 of 12 fail before),
including that the 300/372 spread really is just text length.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 03:03:32 +05:30
debpalashandClaude Opus 4.8 246582341b feat(voices): pick voice-gallery (archetype) voices anywhere via materialize-on-select (#1219)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 05:10:49 +05:30
debpalashandClaude Opus 4.8 9d027b32c9 feat(audiobook): cast/voice mapping + markup toolbar + live stats + validation (#1217)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 01:54:13 +05:30
debpalashandClaude Opus 4.8 dd4af735f3 feat(audiobook): stop/cancel generation + live per-chapter progress (#1216)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 01:38:51 +05:30
debpalash 7f4d0ad9df chore(audiobook): correct issue refs to #1208 + changelog entries 2026-07-20 22:16:50 +05:30
debpalashandClaude Opus 4.8 a12a7e7ee1 feat(audiobook): expressive maturity — overrides, emotion, cache opt-out, discoverability (#1210)
Audiobook renders were locked to the model's most deterministic preset (32
steps / 2.0 guidance / model-default temps) with no way to change it, which is
why books sounded flatter than the same voice on the Voice page. Open that up
without changing any default byte-for-byte.

- Production Overrides in the Audiobook tab: position_temperature,
  class_temperature, num_step, guidance_scale, postprocess_output (+ seed),
  reusing the Voice page's panel. Unset reproduces today exactly.
- IndexTTS2 graded emotion (emo_vector / emo_text / emo_alpha) reaches the
  longform path via a typed engine-options object; engines that don't
  understand an option ignore it (no crash across the ~14 backends).
- Cache opt-out ("vary repeated lines") so identical lines can get distinct
  takes; default off keeps the content-addressed replay.
- Markup reference now lists the reaction tags that already work in audiobooks;
  docs/expressive-speech.md corrected so no recipe it names is unreachable.
- Fix AudiobookGenerateBody dropping `language`, so audiobook language
  selection actually reaches the backend.

Every new param is folded into BOTH cache layers (chapter + segment) and the
per-chapter preview, so changing a knob re-renders instead of replaying stale
audio — and an all-default request keeps its old cache key, so existing books
don't re-render. Regression tests: tests/test_audiobook_expressive.py (backward
compat, cache-signature loop, preview/render parity, engine-ignores-unknown,
cache opt-out, emotion reaches engine) + audiobookOverrides.test.jsx.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 21:45:04 +05:30
debpalashandClaude Opus 4.8 aab138ea57 fix(backend): surface the shell's start-failure diagnosis instead of "can't reach the backend" (#1177)
The reporter's string is apiFetch's LAST fallback, reached only when no crash
marker exists AND the shell's lifecycle stage is 'failed' or 'unknown'. The
'failed' half was the bug: `BootstrapStage::Failed { message }` carries the
whole diagnosis — exit code plus a ~30-line stderr tail, or the precise reason
`ensure_venv_ready` refused (Intel Mac, a failed `uv sync`, a blocked GitHub) —
and `backendLifecycleStage()` returned only the stage tag, throwing the message
away. Every backend-start failure mode collapsed into one generic, evidence-free
sentence that was also factually wrong: it is not starting, and it will not
recover on its own.

- backendLifecycleStage() returns `{ stage, message }`; a `failed` stage gets
  its own branch in apiFetch that surfaces the shell's diagnosis, scrubbed.
- BackendStartFailureNotice renders it after the splash is gone, reusing the
  splash's `detectHints` matcher (shared, so the two can't drift) and the
  existing bug-report affordance. No Retry advice for unrecoverable failures.
- Rust retains the last `Failed { message }` past a later stage transition
  (Retry sets Checking, the supervisor sets StartingBackend) so a respawn can't
  erase the first diagnosis; exposed as `last_bootstrap_failure`.
- `bun desktop` prints the exit code and where to look instead of exiting
  silently — the from-source twin of the same class ("builds but won't launch").
- Scrub primitives extracted to utils/scrub.js so the transport layer can scrub
  without a bugReport -> client import cycle.

Non-Tauri deployments are untouched: there is no shell to fail this way, so the
stage stays 'unknown' and #1164's deployment-specific message still stands.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 14:21:46 +05:30
debpalashandClaude Opus 4.8 f3286c5e6e feat(dub): paste a translation from an external source onto existing segments
After transcription the user can paste a translation produced elsewhere
(ChatGPT, DeepL, a human translator) and have it map onto the segments
that already exist — no re-transcription, no timing loss.

Three input shapes are auto-detected: a timestamped .srt/.vtt (cues matched
to segments by time overlap, greedy one-to-one so one long cue can't be
copied onto several rows), numbered lines (`1.` / `2)` / `[3]`, mapped by
number and falling back to order when a model renumbers mid-answer), and
plain lines (positional, blank lines treated as separators rather than
empty translations). Nothing is applied until the preview dialog has shown
every row as before→after with unmatched rows flagged.

Applying goes through `pasteTranslations` in useSegmentEditing, which
mirrors `segmentEditField`'s duties across rows in ONE undo step: write
`text` and `translations[dubLangCode]` in lock-step and clear the stale
machine-translation badges. It never writes `text_original` (the translate
source `handleTranslateAll` reads — overwriting it would poison every later
re-translate) and never touches a language other than the active one.
Changing `text` alone marks those rows stale via the existing per-language
fingerprints, so no new flag is needed.

The new `POST /dub/parse-subtitle-text` is a stateless wrapper over the
existing `services.srt_parser.parse_srt`, so the lenient cue parsing stays
single-sourced instead of being reimplemented in JavaScript.

Also fixes a ReDoS in that parser, reachable today via /dub/import-srt:
`_TIMING_RE` used `^\s*` under re.MULTILINE, so at every line start the
engine consumed all remaining blank lines before failing on the first
digit — quadratic. 20k blank lines already took 1.7s and a 2 MB blank-line
file never returned, pinning the request thread. Horizontal-whitespace-only
classes make the scan linear.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 13:18:51 +05:30
debpalashandClaude Fable 5 83f943bead fix: bot-review harvest (16 findings) + deterministic style/locale CI + reviewer configs
Harvested and verified every CodeRabbit/Greptile finding from PRs #1175,
#1189, #1192, #1195: 16 real ones fixed (fallback ASR preflight bypass,
VRAM release on stream exit, typed 409 parity, uv env independence,
path-privacy in errors, MCP clone_voice hardening, CaptureWidget WS
guard, test hygiene), 4 refuted with evidence, rest documented as
deliberate design or deferred.

Deterministic CI replaces hand-enforcement: tests/test_changelog_style.py
(quiet one-liner format) and tests/test_locale_parity.py (21-locale
key/placeholder lockstep with a ratchet baseline) — the latter surfaced
and fixes 151 already-broken locale strings. CodeRabbit/Greptile carry
the house rules via .coderabbit.yaml + greptile.json; CLAUDE.md gains
the harvest-before-merge and never-accept-as-is rules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 09:47:41 +05:30
debpalashandClaude Fable 5 8d05507ef8 review: harden dismissed-filter against level escalation; quiet changelog entry with credit
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 09:03:55 +05:30
Ævar Guðmundsson f7da771256 feat(ui): dismissible info/warn system notifications
System notifications are re-generated on every /system/notifications poll,
so standing facts about a machine (gpu-unavailable on a CPU-only box,
disk-low) sit in the bell and footer forever with no way to acknowledge
them. This adds client-side dismissals:

- prefsSlice gains dismissedNotificationIds (persisted, deduped, capped
  at 50) + dismissNotification(id)
- useVisibleNotifications wraps the shared poll with the filter so the
  header bell badge and the footer tab always agree
- info/warn notes render a dismiss button (reuses common.dismiss, no new
  strings); error notes stay undismissible -- broken things remain visible
  until actually fixed
- id stability is the re-arm contract: stable ids stay dismissed across
  sessions, occurrence-stamped ids (last-run-crash-<detected_at>) come
  back as fresh ids exactly as their generator intended
- acting on the run-sentinel unclean-shutdown notice now POSTs
  /system/last-run-crash/ack, mirroring what crash-last-session already
  did -- previously nothing in the UI ever acknowledged it

Tests: store contract (dedupe/cap/level-gate) + render-level coverage of
the footer tab through the real hook chain (mocked API, real QueryClient,
real store); existing suite green (1388 passed), oxlint + tsc clean.
2026-07-19 02:29:50 +00:00
debpalashandClaude Fable 5 63fd497caf feat: TTS-only first run, platform-curated ASR, guided OS permissions, parakeet-mlx
Only the TTS model (~2.4 GB) is required on first run; ASR models are
per-platform curated picks (curated_on in models.yaml) installed on demand.
Every transcription surface returns a typed asr_model_missing error with a
one-click download CTA instead of silently pulling multi-GB Whisper weights.
Settings -> Models is a grouped, platform-aware catalog. New guided
permissions UX (wizard System Check + Settings -> Permissions + mic
pre-flight) with native mic-state checks and OS settings deep-links. New
parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated
capture preference so multilingual dictation never regresses). Docs:
expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ.
Hardening: preflight fails open for custom model pins, ROCm curation no
longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged
consent key, CaptureWidget setup race fixed, offline-cache CI simulation
fixes so empty-cache runners stay green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 15:20:54 +05:30
debpalash c267a632b5 feat(frontend): last-contact tracking + mode-aware honest unreachable-backend errors (#1164)
The give-up error for a transport failure was one-size desktop copy —
'restart the app (or check Settings → Logs → Backend)' — which is
meaningless advice in the two deployments that have no shell: a
'bun run dev' browser session and a Docker/LAN 'server' page. And it never
said the one thing that splits this bug class in half: did the backend
ever answer this session?

- utils/backendContact.ts: apiFetch records a contact timestamp on EVERY
  response (success and HTTP error alike — both prove the process is
  alive; module var + sessionStorage so a reload mid-outage remembers).
  describeLastContact() then tells one of two honest stories: 'it was
  answering Xs ago and then stopped — most likely crashed or was killed
  mid-request' vs 'it has not answered at all this session — it may never
  have started'.
- utils/deploymentMode.ts: desktop (Tauri) / dev (Vite) / server (page
  served by the backend itself — Docker, LAN share, remote GPU).
- api/client.ts: the generic give-up branch now builds the message from
  mode + last contact. dev → points at the bun-run-dev terminal and
  omnivoice.log; server → docker logs/journalctl, noting the page itself
  can go down with the backend when Docker serves it. Desktop keeps its
  existing copy (crash markers + lifecycle stages already cover it).
  Every status-0 ApiError now carries structured diagnostics in .detail:
  {transport, mode, lastContactMs, firstFailureTs, attempts} for the
  bug-report prefill (no consumer read the old string detail).
- i18n: new backendUnreachable.* keys in en.json, mirrored with faithful
  translations into all 20 other locales. The messages resolve through
  i18next when initialized and fall back to self-interpolated English
  when not (node:test loads client.ts without the app bootstrap) — the
  diagnostics path must never crash on the localization layer.

Tests: src/test/client.unreachable.test.ts (12 new) — mode detection,
both last-contact stories end-to-end through apiFetch, contact recorded
on HTTP-error responses, and the ApiError.detail contract; client.test.ts
updated for the structured detail.
2026-07-16 19:18:32 +05:30
Paolo Antinori 55c852c6f2 fix(remote-auth): show an API-key gate (not PIN) for API-key 401s in remote-backend mode (#1154)
* fix(remote-auth): route API-key 401 to an API-key gate, not the PIN form

When OMNIVOICE_API_KEY is set (remote-backend mode), a non-loopback browser
gets 401 "API key required" from BearerKeyMiddleware. But client.ts fired
`ov:pin-required` on every 401, surfacing the PIN gate — whose payload
(sessionStorage ov_pin / X-OmniVoice-Pin) can never satisfy the API-key
middleware. A remote user was stuck on a PIN form they could not pass.

Read the 401 `detail` and dispatch a single `ov:auth-required` CustomEvent
carrying the mode; RemoteAuthGate renders the matching PIN or API-key form.
Adds a `?api_key=` deep-link bootstrap (one-shot — scrubbed from the URL so a
reload can't re-clobber a corrected key) and a guarded saveApiKey helper.

Backend is unchanged — the two 401s are distinguishable by their `detail`
body ("API key required" vs "PIN required"). Docs: remote-gpu.md gains a
"From a browser" subsection for the new ?api_key= deep link.

* fix(remote-auth): preserve URL hash when scrubbing credentials

The replaceState that scrubs ?api_key=/?pin= rebuilt the URL from pathname
(+ optional query) and dropped url.hash, nuking any deep-link fragment
(e.g. #settings). Rebuild with pathname + (?query) + hash.

Addresses greptile + coderabbit review feedback on #1154.

* fix(remote-auth): guard 401 routing against a non-string/malformed detail

String(detail) can itself throw on a 401 detail whose toString is broken
(e.g. { toString: null }), aborting the auth-event dispatch. Match only real
strings with typeof; anything else falls back to PIN mode.

Addresses coderabbit's 17:03 re-review finding on #1154.

* fix(remote-auth): read the deep-link API key from the URL fragment (#api_key=)

Move the remote-backend deep link from ?api_key= (query) to #api_key=
(fragment): fragments are never sent to the server, so the durable key stays
out of the GPU box's and any reverse proxy's request logs on the page load
(greptile P1). ?pin= stays on the query (QR flow, session PIN).

The bootstrap is extracted into a pure, unit-tested _parseDeepLinkCredentials
helper (pin from the query, api_key from the fragment, one-shot scrub of both,
plus a legacy ?api_key= scrubbed-without-reading so a stray query key never
lingers). Docs document #api_key= with encoding guidance for keys containing
+ / & / # / =.
2026-07-16 00:59:06 +05:30
018cdcb47f fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store (#1149)
* fix(dub): hydrate partial translations on tab switch; dialect guard moves into the store

Review round on #1148, both findings real:

- Greptile P1 "missing translations leave mixed text": the in-browser
  translations map can be PARTIAL (tracks generated before per-language
  persistence, partial regens); the non-destructive switch then left those
  rows in the previous language under a single-language preview. New
  GET /dub/segments-text/{job}?lang= exposes segments_i18n (the
  authoritative per-language map every generate rebuilds); the tab click
  hydrates only the gap rows, failure-silent, and skips stale responses if
  the user switched again mid-fetch.
- CodeRabbit "clear stale dialect": the dropdown paths each cleared a
  non-matching dubDialect by hand; the guard now lives inside
  switchDubLangCode so every caller (dropdown, multi-language loop, preview
  tabs, future ones) inherits it. Matching dialects survive.

Tests: endpoint (i18n map served, never-generated track -> empty map, legacy
job -> empty map), hydration (stored rows swap instantly, missing row
hydrates from the mock backend and is cached into translations), dialect
guard (cleared on mismatch, kept on match). Suites: dub sweep 262, frontend
1253, both green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(api): register /dub/segments-text in the route-inventory snapshot

The inventory guard caught the new endpoint exactly as designed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 02:11:43 +05:30
Palash Debnathandmergetest 810b62e62d fix(api): an alive-but-unresponsive backend now says so, instead of "it stopped" (#1113) (#1117)
A v0.3.21 user hit "Can't reach the local OmniVoice backend — it may still be
starting up, or it stopped" — on the release that was supposed to end that class.

Reading the report tells us what happened WITHOUT reproducing it: they got the
generic message, not the crash story. On 0.3.21 apiFetch consults the crash
marker, and a real process death always writes one. No marker ⇒ the backend did
not die. And the shell was still reporting `ready` ⇒ the process was alive.

So both halves of that sentence were false: it had not stopped, and it was not
starting. It was ALIVE and not answering — a job wedged holding the engine
(troubleshooting §14: a generate/transcribe too heavy for the available memory
starves the worker). Telling that user to "restart the app" is the wrong advice
for a stuck job, and it buries the real cause.

When the reconcile window expires and the shell STILL says `ready`, we now know
the process is running, so say that: name the wedged-job cause, point at
Settings → Logs → Backend for what it was last doing, and at a smaller
model/engine as the usual fix. The genuine "stopped or starting" message stays
for the case where the shell has no idea (no shell — browser/Docker), and the
crash story still wins whenever a marker exists.

This does not claim to stop the wedge — it stops the app from lying about it,
and gives the next reporter the right words.

2 new tests; frontend suite 1213 passed.

Refs #1113

Co-authored-by: mergetest <nizam4103@gmail.com>
2026-07-12 17:49:19 +05:30
cbcea41fb6 feat(memory): honest /model/loaded accounting + a free-memory budget probe (#1111)
Two gaps the model-management investigation surfaced, now closed.

1. /model/loaded reported only the OmniVoice core, so a resident second engine
   (mlx-audio, cosyvoice, …) and the warm dictation ASR were INVISIBLE — the
   memory picture looked ~2 GB lighter than reality on exactly the boxes that
   OOM. list_loaded() now enumerates the in-process engine instances (from the
   generate path's cache) and the capture ASR singleton too, and adds a
   `system` block: free/total RAM (and free VRAM on a dedicated GPU) plus a
   low-memory advisory. Verified live: after an mlx-audio generate the panel
   shows `engine:mlx-audio` and `system: {ram_available_gb, ram_total_gb}`,
   where before it showed nothing.

2. services/memory_budget.py: available_memory() reads FREE memory now (device
   caps only reports total, once per process) — free system RAM via psutil,
   free VRAM via torch.cuda.mem_get_info on a dedicated GPU; on MPS the RAM
   figure is what matters (unified memory). low_memory_warning() returns an
   advisory below a headroom threshold (OMNIVOICE_LOW_MEMORY_HEADROOM_GB,
   default 2). The generate path calls log_if_low() before a load, so a later
   OOM kill leaves a breadcrumb pointing at the load that tipped it instead of
   a silent death.

Advisory only — nothing is blocked: the OS reclaims cache, and refusing a load
on an estimate would brick machines that would cope. The single-active-engine
eviction (#1105) is what actually reclaims room; this makes the picture honest
and leaves forensics.

6 new unit tests (threshold logic / VRAM-precedence / never-raises); frontend
LoadedModelsResponse typed for the new `system` field + id shapes. Backend
suite 2918 passed; typecheck clean.

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 16:11:42 +05:30
662339b50a fix(api): don't believe a stale "ready" — close the #1094 race that #1101 hit (#1102)
A 0.3.19 user still got "Can't reach the local OmniVoice backend" (#1101), on
the very release that was supposed to end that class. The fix had a hole.

apiFetch asks the shell whether a start/restart is in progress before erroring.
But the shell's stage is a 2-SECOND POLL, not a live probe: when the backend dies
mid-generate, supervise_backend needs up to ~2 s to notice the child exit, write
the crash marker, and set_stage(StartingBackend). apiFetch asked exactly ONCE, at
the end of the ~2.9 s transport cascade — so it very often still read `Ready` and
fell straight through to the generic toast. Worse, the crash marker usually
wasn't written yet either, so even the honest crash story (#941) was missed and
the user got the vague message.

The bug was trusting `ready` as authoritative. A transport failure CONTRADICTS
it: if the backend were reachable, the fetch would have succeeded. So `ready` is
now treated as a STALE belief — we keep retrying across a bounded reconciliation
window (12 s), re-asking each time, which lets the supervisor catch up and flip
to `starting` (→ the long wait + the restarting banner) and gives the crash
marker time to land so the error can name the real cause. `failed` (the shell
gave up) and `unknown` (no shell — browser/Docker) still error immediately, so a
genuinely dead backend is as prompt as before.

The reporter's trace is the signature: generate:start (design) →
generate:stream-fallback → the toast, on a 16 GB M1/MPS box — i.e. the backend
process died under memory pressure during generation, which is the underlying
crash this now surfaces honestly instead of guessing at.

Regression test reproduces it against the shipped 0.3.19 logic (fails) and passes
on the fix. Frontend 1184 passed; backend 2897 passed.

Refs #1101

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 13:57:10 +05:30