d91beef0fd314250d8d9b94de86dfea019a8bd96
42
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
497d57ee62 |
Show complete engine disk costs before install (#1728)
Closes #1718. Adds structured pre-install and measured post-install disk costs, complete sidecar preflight accounting, strict authorization for recursive scans, and localized catalogue details with confidence and deduplication context. |
||
|
|
0ec9c074e8 | fix: distinguish opaque loaded sidecars | ||
|
|
5a615d2c66 |
feat(workers): package headless GPU nodes (#1638) (#1648)
Closes #1638.\n\nPackages headless GPU workers with durable enrollment, bounded artifact handling, cross-platform lifecycle cleanup, and regression coverage. Incorporates CodeRabbit, Greptile, CodeQL, and platform-CI findings before merge. |
||
|
|
43f1d46fe6 |
fix(indextts): accept the config name upstream ships, and keep long text alive (#1619)
* fix(indextts): accept the config name upstream ships, and keep long text alive Two independent defects, both reported on a working IndexTTS 2.5 install. Install always failed. IndexTeam/IndexTTS-2.5 ships the model config as config.yaml — at the pinned revision d0aa86e7 and at HEAD; config_v2_5.yaml exists in no upstream revision. VoiceStudio demanded that name, so _weights_floor_ok never found it and the install died claiming 'the download was likely interrupted' when the download had been perfect. The only way through was to hand-rename the file. Both names are accepted now, in the installer and on the load path, so installs created with the workaround keep working without a reinstall. Long text was killed at 60s. infer() is one blocking upstream call that puts nothing on the wire, and IndexTTS was the only sidecar still on the 60s recv_timeout_s class default while pockettts and omnivoice-subprocess had both raised theirs. Raising the default alone does not fix it — which is why the reporter's RECV_TIMEOUT_S=3600 edit didn't help: progress frames are also what report activity to the GPU pool's execution clock (#1367), so a silent sidecar still trips the outer generate budget. The sidecar now heartbeats every 5s while infer() runs (and during the cold model construction), _send takes a lock so the beat thread can't interleave framing, and the deadline rises to 900s via OMNIVOICE_INDEXTTS_RECV_TIMEOUT_S. test_indextts25_health_requires_25_config_name asserted the bug — that a checkout holding only config.yaml is unhealthy — so it is rewritten to the corrected contract, including that a genuinely truncated download is still caught. Fixes #1611 * test(indextts): follow the installed config name in the sidecar loader tests Two more tests encoded the config_v2_5.yaml assumption, both asserting cfg_path against a directory where no config existed at all — so they were pinning the literal name rather than the resolution. They now lay down a real checkpoints/ tree and assert the resolved path, including that a checkout carrying the pre-fix hand-renamed config still resolves. Caught by the full suite; the targeted runs during development did not reach tests/backend/services/. * test(indextts): event-driven heartbeat tests, real interleaving proof, precedence pin Review round on #1619 — all four findings taken. - The docs line naming 0.5.1 is version-neutral now ('Earlier installs') — version labels are the owner's call. - The heartbeat tests waited on wall-clock sleeps; they now block on a per-write Event with a bounded deadline, so scheduler load can't flake them. - The _send test asserted the lock EXISTS — a tautology. It now drives four concurrent writers through a stream that yields between every byte and asserts every frame decodes; verified fail-before by removing the lock (torn frame) and pass-after. - The precedence test deleted config.yaml before creating the renamed one, so reversed precedence still passed. Both files now coexist for the assertion; verified fail-before by reversing _CFG_NAMES. |
||
|
|
41722afe3b |
refactor(launchpad): quieter, borderless design refresh (#1515)
* refactor(launchpad): quieter, borderless design refresh The launchpad carried decoration from an earlier direction: icon chips, corner-hung count badges, a permanently visible filled arrow, uppercase mono card titles, and a dotted stipple divider — plus a frame that had been invisible since the app-wide border tokens were zeroed. Rework it around what the borderless direction actually implies: - Feature tiles get a whisper-faint surface instead of a dead frame, and read as three bands (bare glyph + count / title + arrow / description). `--card-hue` is spent sparingly — the glyph at rest, the surface, count and arrow only once raised. Titles move to sans sentence case; counts are plain tabular numerals. Lift softened 4px -> 2px, coloured glow -> neutral shadow, plus an explicit focus ring and a staggered entrance. - Hero drops the boxed "646" pill and the filled A/B-Compare button for quiet type, with a hairline standing in for the separation. - Section labels trade the dotted stipple for a single fading hairline; rows are transparent until hover and reveal "Open" on hover/focus (it stays in the DOM, so AT and keyboard always reach it). - Hero, tiles, recent files, callout and project lists now share one 1180px column — previously only the top half was capped, so lists ran edge-to-edge on a wide display while the deck stayed centred. Two bugs found and fixed while doing it: - Buttons that had `border border-solid border-transparent` removed fell back to the UA default border and rendered a visible 1px outline. They now carry `border-0` explicitly. - `.lp-animate` used `animation-fill-mode: both`, so after the entrance it kept owning `transform` — and animation-origin declarations outrank normal ones, which silently killed the card hover lift. Now `backwards`, which still holds the from-state through the stagger delay. Also drops CSS the page has not rendered since #904: the cursor-spotlight layer, the breath ring, and the per-card waveform strip. Verified with headless renders at 1600/1280/940 and the empty state. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(dictation): decode Wayland portal signals and show the capture pill The GlobalShortcuts portal declares Activated/Deactivated as (o session, s shortcut_id, t timestamp, a{sv} options). We decoded the timestamp as u32, so zbus rejected every signal with Signature mismatch: got `(osta{sv})`, expected `(osua{sv})` and the press was dropped as an invalid signal. Registration succeeded and the desktop even reported the bound chord back, so the hotkey looked wired up while doing nothing at all — on every Wayland compositor, for the whole life of the feature (#1490). Decode the 64-bit timestamp, and keep the 32-bit spelling as a fallback so a non-conforming portal degrades to working rather than to silence. With presses arriving, the second half of the failure showed: nothing had shown the widget window since it became a hidden recorder host, so a capture ran with no pill on screen — and a mic or Accessibility failure rendered into a window nobody could see. Add show_dictation_pill, which bottom-centres the capsule on the monitor under the pointer and shows it without taking focus (Windows keeps SW_SHOWNOACTIVATE so paste still lands in the user's document), and call it from the widget for every state but idle. Wayland denies clients their own placement, so the compositor picks the spot there; the pill still appears. dispatch_dictation_capture now logs whether a press was emitted or queued — a press that reaches Rust and produces nothing was otherwise indistinguishable from one the compositor never delivered. Tests: portal signals decode at both timestamp widths (the 64-bit case fails before this change with the exact production error); pill placement centres, respects a second monitor's origin, and clamps rather than going off-screen; the widget shows for a state needing the user, stays hidden while idle, and never shows for a press that arrives while dictation is disabled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: sync in-progress workspace changes Uncommitted work already in the tree, checkpointed so the branch matches the local machine: - Remote GPU workers: join-from-the-app flow, one-time secrets, QR join codes, a Compute control in the status bar, and the device-list Workers panel (#1516) - Model Catalogue workspace, with Settings pointing at it - Settings sidebar search and keyboard navigation - Demo assets for dubbing, dictation and voice design, plus the scripts that render them - Backend: validation-error handling, ASR request-path degradation, and the accompanying tests - CHANGELOG entries for the above and for the Wayland dictation fix Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(tests): follow Engines to the Model Catalogue, and green the sweep - test_supertonic3 asserted the license gate points at "Settings" while the engine now names Model Catalogue → Engines, which is where the accept button actually lives. The assertion follows the move; what it pins is unchanged — the hint must name a place the user can reach it. - Carries the CJK allowlist entries for the rendered dub bundle (#1517) and the regenerated route snapshot for /workers/agent (#1516), both of which this branch inherits from the workspace sync. - docs/install/linux.md: the dictation capsule is bottom-anchored everywhere except Wayland, where the protocol gives applications no say in their placement. Documented rather than left as a surprise (CodeRabbit). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: stop a flaky dependency fetch from failing green runs en-core-web-sm resolves to a direct GitHub release URL, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own three retries all land within the same few seconds and fail together, so the whole job dies on a dependency that has nothing to do with the change under test — it cost #1518 and #1517 an otherwise-green run tonight. Two changes: back off between whole `uv sync` attempts, which is what actually clears it, and pass --no-sync to the pytest steps. `uv run` re-resolves the environment before running, so every test step was a fresh chance to hit the same fetch even though the install step had already synced — that is exactly how #1518 failed, in the isolated backend/tests step, with all 5467 tests already passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: one retry seam for every uv sync, not just the job that failed last en-core-web-sm resolves to a direct GitHub *release* URL rather than a package index, and github.com intermittently answers `http2 error: refused stream before processing any application logic`. uv's own retries all land inside the same ~10 seconds and fail together, so a job dies on a dependency unrelated to the change under test. Tonight that cost four otherwise-green runs across #1515, #1517 and #1518 — and the first fix only covered the Tests job, so the next failure simply moved to Smoke (Linux), which syncs separately. The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh backs off between whole attempts (15s, 45s, 90s) and every workflow that syncs now goes through it — ci.yml (tests + the platform matrix), release.yml, security.yml, evals.yml. It still fails loudly after four attempts, so a genuinely broken lockfile is not disguised as a flake. The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke matrix has always set, which is part of why it was the one that kept dying; it has them now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ci): pin the Intel-Mac contract by intent, not by command spelling test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted the literal line `run: uv sync --extra pockettts`, so routing every sync through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The contract it exists to protect is that the pockettts extra installs ONLY on backend_supported legs — which the regex now pins, while leaving how the sync is invoked free to change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci: keep every uv run out of the resolver, and bound the retry budget CodeRabbit, #1517: - `uv run` re-resolves before running, so the smoke suite, the worker-artifact tests, the release test run and the eval run were each a fresh chance to hit the flaky direct-URL fetch outside the retry loop. All of them pass --no-sync now; the environment is already synced by the step that owns the retries. security.yml's `uv run --with pip-audit` is deliberately left alone — it layers an ephemeral package rather than running the project's own tests. - The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a 120 s timeout on the smoke matrix). Three attempts and 60 s of total backoff outlast the refusals actually observed while staying well inside the jobs' timeout-minutes. - The Intel-Mac contract test pinned the smoke command literally too, so --no-sync tripped it exactly like the sync line did. Same fix: assert the contract (smoke runs only on backend_supported legs), not its spelling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
95a35b8e07 |
feat(indextts): add native IndexTTS 2.5 support (#1485)
* feat(indextts): add native 2.5 sidecar support * fix(indextts): preserve legacy language metadata * docs(indextts): state model license terms accurately * fix: preserve IndexTTS upgrades and duration controls * fix: complete IndexTTS upgrade safeguards |
||
|
|
a2745f1029 | test: assert fixed-shape secret log record | ||
|
|
32d9a8a964 | fix: tighten log safety regressions | ||
|
|
e086cb03f1 |
Merge remote-tracking branch 'origin/main' into fix/ghas-log-safety
# Conflicts: # CHANGELOG.md |
||
|
|
38a00cbf30 |
fix(security): stabilize engine discovery metadata (#1460)
* fix(security): stabilize engine discovery metadata * fix(security): preserve stable routing outcomes * fix: preserve safe engine routing outcomes |
||
|
|
fe856c9c5e | fix(security): omit sensitive log context | ||
|
|
bcb547b9f2 |
fix(engines): a slow venv probe is not a broken venv (#1414) (#1421)
Every subprocess engine confirms a candidate interpreter by spawning it and importing the engine package. For IndexTTS that is 'import indextts.infer_v2', which pulls in torch and transformers — seconds with a warm page cache, tens of seconds on a first run, a spinning disk, a network share, or Windows with real-time AV scanning every DLL. The bound was 10s (15s for three peers), and elapsing it was treated as a negative: the candidate was discarded exactly as if the import had raised. A working OMNIVOICE_INDEXTTS_DIR install was reported as 'IndexTTS-2 is not installed', or fell through into the lazy bootstrap and reinstalled over a working clone. Only successful resolution was memoised, so every retry re-ran the probe and failed identically — which is why all three reported repro paths look like one bug. A timeout is the absence of evidence, not evidence of breakage. The probe is now tri-state: yes (imported), no (ran and failed), unproven (did not finish). An unproven candidate is kept as a fallback and used only after every candidate has had its chance, so a wedged user clone cannot shadow a healthy bootstrapped venv. If an unproven venv really is broken it now fails at the sidecar handshake with a real error rather than a confident lie about the install. Fixed as a class: backend/engines/_venv_probe.py replaces the drifted copy in each of the four bootstraps, and the bound is tunable per engine, defaulting to 60s. Zero and negative values are ignored — an unbounded probe would let one wedged candidate hang engine resolution forever. Reported with a precise root cause by @OracleNightmare. (#1414) |
||
|
|
6cfef5c0cf |
fix(routing): warn about an under-provisioned GPU before the job, not after (#1226, #1222)
Two users on 4 GB cards (GTX 1650 Ti, Quadro P2000) ran the `omnivoice` engine, waited out the full compute budget, and were told the job "was too heavy for the available compute … most often the GPU is VRAM-starved". The 300s-vs-372s spread between the two reports is purely text length (`300 + (len-1200)/40`, so 372s ⇒ ~4080 chars) — one bug, not two. Nothing about the budget is device-aware, and nothing needs to be: the real defect is that until the moment it failed, routing showed a clean green "accelerated". `resolve_routing` matched on GPU *family* only, so a 4 GB card and a 24 GB card were indistinguishable, and no engine declared a VRAM requirement anywhere in the repo. - `TTSBackend.min_vram_gb` — advisory metadata alongside `gpu_compat`. Only `omnivoice` declares one (6 GB), derived from the pool's own measured per-job budget (`_GPU_VRAM_PER_JOB_GB = 5.0`) plus resident weights. Inventing floors for engines with no measured figure would put confident numbers in the UI that nothing backs. - `resolve_routing` takes the floor and emits an accelerated-with-caveat reason when the host is below it. Reuses the existing caveat channel, so the Settings matrix and the synth-time routing notice surface it with no UI change. Advisory, never blocking: drivers page to system RAM, and short inputs fit where long ones don't. Kernel-risk still outranks it, and a failed VRAM probe (0.0) never guesses. - `_timeout_guidance` names the actual card and its VRAM, and leads with "pick a lighter engine" instead of wording that reads as transient contention the user can flush their way out of. Regression test: tests/test_low_vram_advisory.py (8 of 12 fail before), including that the 300/372 spread really is just text length. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7f4d0ad9df | chore(audiobook): correct issue refs to #1208 + changelog entries | ||
|
|
a12a7e7ee1 |
feat(audiobook): expressive maturity — overrides, emotion, cache opt-out, discoverability (#1210)
Audiobook renders were locked to the model's most deterministic preset (32
steps / 2.0 guidance / model-default temps) with no way to change it, which is
why books sounded flatter than the same voice on the Voice page. Open that up
without changing any default byte-for-byte.
- Production Overrides in the Audiobook tab: position_temperature,
class_temperature, num_step, guidance_scale, postprocess_output (+ seed),
reusing the Voice page's panel. Unset reproduces today exactly.
- IndexTTS2 graded emotion (emo_vector / emo_text / emo_alpha) reaches the
longform path via a typed engine-options object; engines that don't
understand an option ignore it (no crash across the ~14 backends).
- Cache opt-out ("vary repeated lines") so identical lines can get distinct
takes; default off keeps the content-addressed replay.
- Markup reference now lists the reaction tags that already work in audiobooks;
docs/expressive-speech.md corrected so no recipe it names is unreachable.
- Fix AudiobookGenerateBody dropping `language`, so audiobook language
selection actually reaches the backend.
Every new param is folded into BOTH cache layers (chapter + segment) and the
per-chapter preview, so changing a knob re-renders instead of replaying stale
audio — and an all-default request keeps its old cache key, so existing books
don't re-render. Regression tests: tests/test_audiobook_expressive.py (backward
compat, cache-signature loop, preview/render parity, engine-ignores-unknown,
cache opt-out, emotion reaches engine) + audiobookOverrides.test.jsx.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
83f943bead |
fix: bot-review harvest (16 findings) + deterministic style/locale CI + reviewer configs
Harvested and verified every CodeRabbit/Greptile finding from PRs #1175, #1189, #1192, #1195: 16 real ones fixed (fallback ASR preflight bypass, VRAM release on stream exit, typed 409 parity, uv env independence, path-privacy in errors, MCP clone_voice hardening, CaptureWidget WS guard, test hygiene), 4 refuted with evidence, rest documented as deliberate design or deferred. Deterministic CI replaces hand-enforcement: tests/test_changelog_style.py (quiet one-liner format) and tests/test_locale_parity.py (21-locale key/placeholder lockstep with a ratchet baseline) — the latter surfaced and fixes 151 already-broken locale strings. CodeRabbit/Greptile carry the house rules via .coderabbit.yaml + greptile.json; CLAUDE.md gains the harvest-before-merge and never-accept-as-is rules. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
933743e336 |
fix: resolve open issue batch — #1172 #1173 #1174 #1185 #1186 #1188
- validate managed binaries before exec; 0-byte GGUF placeholders fail actionably instead of "Exec format error" (#1172) - KittenTTS: tokenizer-measured chunking to the ONNX 512-token cap; clear 400 for unspeakable input (#1173) - clean SIGTERM during weight load: shutdown-aware loader, benign cancelled-load classification, lifespan hardening, scoped log silencers (transformers load + alembic fileConfig) (#1174) - broken ASR deep-imports (lightning_fabric) mark the engine unavailable with a repair hint and fall through (#1185) - uv cache + managed Python follow the chosen install drive on Windows (cherry-picked cross-drive class fix + spaces/D: tests) (#1186) - adaptive silence-removal ladder for quiet clone references; localized actionable error for truly silent clips, all 21 locales (#1188) - CHANGELOG: consolidated Unreleased into the quiet one-liner style Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ff56865cf7 |
feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines (#1083)
* feat(engines): one-click IndexTTS-2 sidecar install from Settings → Engines
IndexTTS-2 required four manual terminal steps (git clone, uv venv,
uv pip install -e ., export OMNIVOICE_INDEXTTS_DIR). This turns that into
a guided in-app install:
- backend/services/sidecar_install.py — parametrized sidecar provisioner
(SidecarSpec/SPECS so future sidecar engines are one entry, not another
installer). Resumable background job with step-by-step status: disk-space
preflight (needs-X/have-Y message), source fetch (git clone --depth 1
primary, GitHub tarball fallback when git is absent/fails), dedicated
venv via uv (OMNIVOICE_BUNDLED_UV → PATH resolution; transformers<5
isolation preserved — the parent env is never touched), import-probe
verification, IndexTeam/IndexTTS-2 weights into <checkout>/checkpoints
(where the sidecar actually loads from) via snapshot_download with the
auto-selected/configured HF endpoint + token — no hardcoded
huggingface.co — and persistence of OMNIVOICE_INDEXTTS_DIR (os.environ
for immediate use, prefs.json env.* for the next launch). Idempotent:
partial installs repair, downloads resume, healthy installs (incl. a
user's own clone) report already_installed and are never touched.
- API: POST /engines/{id}/install starts the job, GET
/engines/{id}/install/status polls it, DELETE /engines/{id}/install
removes an app-managed install (loopback-gated; refuses user-managed
clones). list_backends() gains one_click_install.
- Frontend: Settings → Engines shows an Install button on the IndexTTS2
row with per-step progress, live log tail, weight-download %, and
error+remediation; the manual setup snippet is demoted to a collapsed
"Manual install" fallback. All strings via i18n (en.json).
- OMNIVOICE_INDEXTTS_DIR joins the Settings env-var allowlist
(single-sourced from the installer SPECS).
- Docs: docs/engines/indextts.md leads with the one-click flow; manual
steps become the fallback section. CHANGELOG Unreleased entry added.
- Tests: tests/test_sidecar_install.py (24 cases — happy path, disk-space
fail, git-absent/git-failing tarball fallback, partial-install repair,
already-installed/running gating, uninstall safety, spec↔bootstrap
contract, router wiring) + 6 new EngineCompatibilityMatrix RTL cases.
API route snapshot regenerated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): harden the sidecar installer — review findings
- Route namespace: /engines/sidecar/{id}/install — a dynamic
/engines/{id}/install would shadow the literal
POST /engines/sonitranslate/install (engines router registers first);
regression-guarded by test_sidecar_routes_never_shadow_literal_engine_routes.
- Weights completion marker: a killed-mid-download multi-shard weights dir
(config.yaml + plausible shards) no longer passes for healthy; the marker
is written only after snapshot_download returns, so re-runs resume.
- _run_logged: drain thread + proc.wait(timeout) + POSIX process-group kill
— a grandchild holding the stdout pipe can no longer hang the step past
its timeout.
- Job log lock: the status poll's list(deque) copy no longer races the
worker's appends (RuntimeError under active logging).
- Self-heal: a healthy managed install whose env var was lost (prefs wiped)
is re-pointed by start_install instead of reported already_installed
while the engine stays unavailable; legacy bootstrap installs (Probe-2
venv) are trusted via the engine's own probe.
- Single-sourced uv/venv-layout resolution: engines.indextts.bootstrap now
delegates _locate_uv/_venv_python_path to services.sidecar_install.
- Frontend: stable poll interval (keyed on the running-id set, not the
status map), reload on a job that finishes before the first poll,
re-attach to an in-flight job on remount, i18n'd Install aria-label,
manual-install <details> auto-opens on failure, snippet block hoisted
out of the JSX IIFE.
- list_backends: sidecar-installable set hoisted out of the per-engine
loop; exhaustive-shape registry test updated for one_click_install.
- Tests rebind the live services.sidecar_install module per test (other
suites purge sys.modules["services"], which made router tests
order-dependent).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): fill in the PR ref (#1083)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(security): validated tarball fallback + scanner-clean installer
- The pre-filter= extractall fallback (Python < 3.11.4) now extracts
member-by-member behind the same guards extractall(filter="data")
enforces — regular files/dirs only, no absolute paths, no ../ escapes,
resolved-path containment. Kills the new CodeQL py/tarslip (high) and
Bandit B202 (error) alerts; regression-tested with a malicious tarball
(test_safe_extract_members_blocks_tar_slip).
- snapshot_download tracks the weights repo's default branch on purpose
(same policy as every other model download; artifacts are
checksum-verified by hf_hub) — documented + B615 waived at the call.
- Explanatory comments on the intentional empty-except blocks
(CodeQL py/empty-except notes).
Verified locally: bandit -ll -ii on the module reports 0 MEDIUM+ findings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(engines): address Greptile review — Windows tree kill, prefs write race, poll robustness
- _kill_tree: Windows now uses taskkill /F /T so a git/uv helper spawned by
the timed-out child can't keep writing into the checkout (POSIX already
killed the process group). Unit-tested with os.name patched to nt.
- core/prefs: mutations (set_/delete) are serialized behind a module lock —
the installer worker persisting its env.* key concurrently with a Settings
write could previously drop whichever key saved first (whole-class fix:
every threaded prefs writer, not just the installer). Fail-before/
pass-after: tests/test_prefs_thread_safety.py.
- Matrix polling: at most one in-flight status request per engine (an old
'running' response can no longer land after a newer 'succeeded' and
restart the poller), and four consecutive poll failures drop the stale
snapshot instead of showing "Installing…" and hammering a dead backend
forever.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
17ae952810 |
feat(settings): Models & Engines pages — engine identity marks, capability badges, upgrade hints, filter, residency (#1058)
The engine list gains a scannable identity mark per engine (EngineMark), capability badges (cloning, device routing with reasons, sidecar isolation), and surfaces available-but-has-advice hints that list_backends previously dropped (new additive hint field; the ready-with-advice convention). The model store gains a filter, disk context near downloads, in-memory residency indicators with safe unload, copyable setup snippets, and actionable empty/error states. Registry additions are additive only (hint, supports_cloning with the property-descriptor guard). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
549fa4009f |
feat(engines): expose MLX-Audio's curated model picker (#981) (#994)
mlx-audio multiplexes 7+ curated models (Kokoro, CSM, Qwen3-TTS, Dia,
Chatterbox, MeloTTS, OuteTTS) behind a single "mlx-audio" backend id, but
MLXAudioBackend resolved its active model ONLY from the
OMNIVOICE_MLX_AUDIO_MODEL env var — invisible to Settings and unreachable
without restarting the packaged app with that var set. A user who
downloaded e.g. Llama-OuteTTS via Settings → Models had no way anywhere
in the UI or API to actually load it; the backend silently kept using
Kokoro.
Fix:
- MLXAudioBackend.__init__ now resolves its model via
prefs.resolve("mlx_audio_model_id", env=..., default=...), mirroring
active_backend_id()'s env > prefs > default order exactly.
- get_active_tts_backend()'s switch-detection now also tracks the
resolved mlx-audio model key, so a model-only change (same backend id)
invalidates the cached instance and reconstructs it — no app restart
needed to pick up a different curated model.
- POST /engines/select gained an optional model_id field; for
family=tts/backend_id=mlx-audio it validates against
MLXAudioBackend.CURATED_MODELS (or a raw HF repo id, matching the
class's existing tolerance) and persists it via prefs.
- GET /engines now includes a curated_models roster + active_model_id on
the mlx-audio entry only.
- Settings → Engines renders a small model dropdown on the mlx-audio row,
pre-selected to the active model, wired through selectEngine's new
optional modelId argument.
Regression coverage: prefs resolution + env override, cache invalidation
on model-only switch, /engines/select 400s on an unknown model id and
persists a valid one, curated_models present only on mlx-audio, and a
new EngineCompatibilityMatrix vitest suite for the dropdown.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
b6ec4e23f3 |
fix(engines): snapshot lazy registry keys so /engines can't 500 under concurrency (#940)
* fix(engines): snapshot lazy registry keys so /engines can't 500 under concurrency
`list_backends()` runs in a FastAPI threadpool and iterates the lazy TTS/ASR
registries via `items()` → `__iter__`, which held a *live* `dict.__iter__(self)`
open across each engine's slow `is_available()` probe. Meanwhile the lazy
`__getitem__` resolves a deferred entry by mutating the dict (`self[key] = cls`).
A second concurrent `/engines` request (or any ASR op) materializing the lazy
`faster-whisper-isolated` entry therefore changed the dict size mid-iteration:
RuntimeError: dictionary changed size during iteration
asr_backend.py:1729 list_backends → _REGISTRY.items()
asr_backend.py:1665 __iter__ → for k in dict.__iter__(self)
Both `_LazyRegistry` (TTS) and `_LazyASRRegistry` (ASR) now snapshot their live
keys up front with `list(dict.__iter__(self))` — consumed atomically under the
GIL — so a concurrent lazy insert can no longer trip the iteration. The slow
per-engine probes then run over the snapshot, not the live iterator.
Deterministic fail-before/pass-after regression for both registries:
tests/backend/services/test_lazy_registry_concurrency.py.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(changelog): add the /engines concurrency fix under [Unreleased] (#940)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
5bd8968aea |
feat(engines): real synthesis "Self-test" + copy-paste setup snippet for opt-in engines (#930)
* feat(engines): real synthesis "Self-test" + copy-paste setup snippet for opt-in engines Builds on #905's Engines-settings fixes (verified still green: license dialog mounts, matrix reloads on select, cpu_fallback routing toast, cpu-native → cpu_only). Two enhancements, no #905 behavior touched. Real "Self-test" for in-process TTS engines ------------------------------------------- The existing /engines/{id}/health probe only imports the package and reports "deps OK" for in-process engines — it never proves the engine can emit audio. New POST /engines/{id}/selftest runs a *tiny real synthesis* from a fixed short ASCII phrase and reports ok + duration + sample-rate + sample count, proving the engine actually produces audio. Guardrails keep it cross-platform-identical and CPU-cheap: TTS + available + in-process only, bounded wall-clock timeout (OMNIVOICE_SELFTEST_TIMEOUT_S, default 90s) that returns ok=false/timed_out instead of hanging the panel, a process-wide lock so a click-storm can't stack model loads, loopback-gated, and only ever on user click (never on load). The Compat Matrix gains a "Self-test" button (with cooldown) that renders "0.82s @ 24 kHz in 820 ms". HF tokens in a synth error are redacted like the health route. Verified end-to-end: kittentts synthesized 89,200 samples @ 24 kHz. Copy-paste setup snippet for path-gated opt-in engines ------------------------------------------------------ IndexTTS / MOSS-v1.5 / dots.tts / Confucius4 gate on an OMNIVOICE_*_DIR env var. list_backends() now emits a single-sourced `setup_snippet` (the exact `export VAR=/path/...` line) surfaced with a Copy button inside the matrix's "Why unavailable?" disclosure, so users don't reconstruct it from the docs. Also tightened the incomplete SelectEngineResponse TS type to include the routing echo (routing_status/effective_device/routing_reason) the post-select toast already reads at runtime. Tests: backend selftest success/subprocess-reject/unavailable/unknown/loopback/ exception-capture/timeout/HF-redaction + setup_snippet shape; frontend self-test render, timeout marker, subprocess+ASR gating, setup-snippet render. New route added to the API route snapshot. Full vitest (808) + backend engine/routing/asr/ route-inventory/no-CJK green; lint 0 errors; format + typecheck:ci clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): allow setup_snippet key in list_backends shape assertion The engine self-test PR added setup_snippet to each backend entry but only updated the route-shape test; test_list_backends_shape strict-asserts the key set. Add setup_snippet there too. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
75864a597f |
fix(dictation): refinement never stalls a final (~51s→≤4s), REST polish parity, real ASR preload reuse (#911)
P0 — Refinement blocked every dictation final with no timeout. With refinement auto:true and a slow/dead LLM endpoint, maybe_refine ran unbounded and blocked the final send in all three capture_ws handlers (~51s measured; the pill hung "Transcribing…" until the widget's 15s fallback fired). Fix the class: a hard, env-tunable budget (OMNIVOICE_REFINE_TIMEOUT_S, default 4s) via a new maybe_refine_async — a slow/dead endpoint now falls back to the unrefined (but polished) text within the budget and can NEVER delay the final beyond it. The LLM HTTP call is bounded to the same budget so the orphaned worker unwinds instead of holding a connection for the client's full 45s. Refinement is now also fully best-effort in the legacy handler (it can't turn a good final into an error frame). P1 — REST /transcribe lacked polish parity. capture.py never applied polish_text, so REST returned raw "…test" while the WS returned "…test." Apply text_polish.polish_text to `text` and `refined_text` (segments stay raw), so the widget POST fallback and MCP/CLI callers match the live socket. P1 — The #888 "instant first dictation" preload was a no-op. The preload called warmup() only `if hasattr`, but SherpaDictationBackend had none, and the WS handlers built a FRESH backend per session so a warm singleton wasn't reused. Add SherpaDictationBackend.warmup() (builds the recognizer) and share one warm recognizer per model id across sessions (get_sherpa_dictation_backend, same invalidation + a shared lock as the capture singleton); each session keeps its own decode stream. First dictation no longer pays the 1.3–2.5s load. P1 — llm_ready is a lie (feeds the P0). It only means "an endpoint is configured", so a placeholder key reads as ready. The P0 timeout makes a dead endpoint harmless; add last_refine_status so RefinementPanel flags a configured-but-failing LLM and links to LLM Providers → Test. Regression tests (fail-before/pass-after): slow-LLM WS final arrives < budget; maybe_refine_async hard timeout + status; REST polish parity + refined_text polish; warmup builds the recognizer and a second session reuses it; the panel honesty note. Backend refinement/capture_ws/capture/sherpa suites, CJK + route inventory gates, full vitest (733), lint (0 errors) and format all green. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
83e71c5689 |
fix(asr): close the #730 residuals — chunked dub wedge shares the guarded reset; repeated timeouts recommend the crash-isolated engine (#895)
Residual A — the chunked dub-stream had a PARALLEL wedge mechanism (its own ping-loop timeout, its own _reset_pool_on_wedge, a dead-end "Try restarting the server" message). A wedged chunk now routes through the SAME run_transcribe_guarded bound+reset as the whole-file paths (#731/#851): the guard resets the poisoned pool once per wedged attempt (no double-reset on retry) and the user sees the actionable ASRTimeoutError. The reset logic is extracted to asr_backend.reset_pool_after_wedge — one shared mechanism, so the semantics can't drift again. run_transcribe_guarded also gains a timeout_env param so chunk errors name OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S instead of the whole-file knob. Residual B — the crash-isolated ASR sidecar (#393, faster-whisper-isolated) is wired as an explicit ESCAPE HATCH, not a default: - selectable end-to-end: Settings engine list gets an explanatory install_hint; honest gpu_compat ("cuda","cpu" — it wraps the same CTranslate2 engine as faster-whisper); get_active_asr_backend now hands back a process-wide singleton for subprocess-isolated backends (a fresh instance per request would leak atexit hooks and respawn the sidecar — reloading its model — on every transcribe). - on the SECOND consecutive guarded timeout-with-reset in one session (resets aren't recovering the hang; the wedged thread keeps its VRAM), the error the user sees + the log recommend switching to the isolated engine in Settings → Engines. Never auto-switched (owner rule: no silent behavior divergence); a completed transcribe resets the streak. Tests (fail-before/pass-after verified against origin/main): wedged-chunk SSE integration (reset count + actionable error + recommendation surfaces), consecutive-timeout streak (fires at 2, resets on success, suppressed when already on the isolated engine), timeout_env parametrization, shared-reset helper, isolated backend in list_backends with hint + honest availability, singleton caching, gpu_compat matrix entry. Docs: troubleshooting §14 gains the chunk knob + escape-hatch guidance. Closes the residuals tracked on #730. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
129beb0ee6 |
test(settings): de-flake the at-rest-encryption assertion (#469)
test_stored_value_is_encrypted_not_plaintext asserted `"hf_" not in raw`, but the stored value is Fernet URL-safe base64 whose alphabet includes `_`, so a random ciphertext occasionally contains the substring `hf_` by chance — a false failure that bit unrelated PRs on CI (~1 in N runs). Replace the 3-char-prefix substring check (weak AND flaky) with stronger, deterministic guarantees: - the full token is absent from the raw column (kept), - a 16-char leading chunk is absent (no partial leak; 62^16 ≈ never collides), - and the value round-trips via get_hf_token() — proving it's genuinely encrypted, not merely absent/empty. Verified non-flaky: the target test passed 8/8 consecutive runs. |
||
|
|
8c8d525397 |
feat(routing): wire effective-device into /engines + select gate (#21 PR 3/5) (#432)
* feat(routing): wire effective-device + routing_status into /engines (#21 PR 3/5) Surfaces the PR-1 probe + resolver through the engine registries so the matrix UI (PR 5) and the no-silent-fallback gates can consume it. - `engine_routing.routing_fields()`: shared helper returning the three serialization-ready keys, centralizing the scrub rule — routing_reason is scrubbed via `core.scrub.scrub_text` only when truthy, so a None reason stays JSON `null` (never coerced to ""). - TTS/ASR `list_backends()` each gain `effective_device` / `routing_status` / `routing_reason`, computed from a SINGLE `detect_host_caps()` call per request (host caps are constant per process). ASR is brought to full TTS parity: it now also carries `install_hint` / `last_error` / `isolation_mode` and a SCRUBBED `reason` (closing a pre-existing ASR token-leak gap) — an identical 11-key shape across families. ASR also gains the same is_available()-raises resilience TTS has (degrade to available:false, never 500). - LLM `list_backends()` reaches 11-key parity too but emits literal `effective_device:"network"` / `routing_status:"n/a"` / `routing_reason:null` (NOT via resolve_routing — LLM runs no local GPU model). `LLMBackend.gpu_compat = ()`. "network" is a label, not a probe — nothing here touches the network. - `select_engine` host-routing gate: refuses a pick whose `routing_status` is `unavailable` on this host (400 with an actionable detail), while ALLOWING `cpu_fallback` (it runs, just slower). LLM is never gated. Defensive `.get` so legacy payloads still select. New typed `SelectEngineResponse`. Tests: 11-key shape across all 3 families, well-formed tts/asr routing keys (+ None-not-"" contract), LLM network/n/a labels, select gate (block unavailable / allow cpu_fallback / never-gate LLM). Updated the registry exact-shape test for the 3 new keys. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(cjk): allowlist docs/specs/ in the hardcoded-CJK guard PR #429 merged the longform design specs, which legitimately quote functional CJK (test-fixture descriptions, CosyVoice speaker IDs, multilingual sample text). The CJK guard scans every tracked file, so those docs turned main red. Specs are documentation, not shipped UI strings — allowlist the docs/specs/ prefix, matching the individually-allowlisted docs already in the set. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c3b2346759 |
fix(engines): MLX platform gate (#390) + ASR gpu_compat + IndexTTS2 (#21 PR 2/5) (#431)
Builds on the device probe from PR 1. Backend-only; the routing keys are wired into /engines in PR 3. - #390 closed: MLXAudioBackend / MLXWhisperBackend now call the shared `core.device_caps.mlx_supported()` gate FIRST, before importing the package. On Linux/Windows/mac-Intel they report unavailable and never advertise a usable `mps` route, even with a stray mlx wheel installed. Replaces the ASR backend's ad-hoc inline MPS check with the one shared rule. (The Wave-4.4 OSError/RuntimeError import-guard is preserved — it now lives behind the platform gate; its test forces the gate open so the guard stays the path under test.) - `ASRBackend` ABC gains `gpu_compat: tuple[str, ...] = ("cpu",)` mirroring TTSBackend, and each subclass declares its real targets: whisperx/faster-whisper → (cuda,cpu); mlx-whisper → (mps,cpu); pytorch-whisper → (cuda,mps,cpu); nemo/funasr → (cuda,cpu); moonshine → (cpu,). Inert until PR 3 serializes them. - IndexTTS2 declares `gpu_compat = ("cuda","cpu")` so it stops advertising the inherited CPU-only default. - ROCm is deliberately NOT claimed for any ASR engine (or for IndexTTS2): CTranslate2 has no upstream HIP build, and an unverified `rocm` claim would route ROCm hosts to a broken GPU path — strictly worse than the honest `cpu_fallback` the resolver already emits ("declares CUDA only; ROCm not in its compat set"). The per-engine TTS ROCm audit is a tracked follow-up that will verify each path before claiming it. Tests: MLX gate regression (both backends, on/off Apple), ASR gpu_compat tuples + no-false-rocm invariant, IndexTTS2 override; existing MLX import-guard test updated for the new gate ordering. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
599f3bcc5c |
feat(engines): on-demand unload of subprocess-engine sidecars (Action 13) (#406)
Completes the dynamic engine load/unload slice. The idle reaper (#401) frees sidecar VRAM after 5 min; this adds a user-initiated "free VRAM now" path so multi-engine users don't have to wait: - subprocess_backend: `list_live_sidecars()`, `unload_sidecar(id)`, `unload_all_sidecars()` via a shared `_force_reap(predicate)` — busy-guarded exactly like the idle reaper (non-blocking lock; a sidecar mid-synth is skipped, never interrupted; next request respawns it). - system.py: `/model/loaded` now surfaces live sidecars as unloadable rows; `/model/unload/{sidecar:<id>|sidecars}` frees one or all. The existing generic flush panel picks these up with zero frontend change. Also refresh CLAUDE.md stale version notes: main is 0.3.6 (latest release v0.3.5 + 1 patch); the v0.3.0-as-unreleased framing in the project/cadence notes is corrected to the v0.3.x continuous-to-main reality. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
34c8ab2409 |
feat(engines): idle-reap subprocess-engine sidecars to free VRAM (Wave 13) (#401)
Parity Action 13 (dynamic load/unload), subprocess-engine half. A subprocess engine's sidecar holds a process — and, for GPU engines, VRAM — for the life of the backend, even after the user switches engines. The default in-process OmniVoice model already idle-unloads (model_manager.idle_worker); this gives the subprocess engine class the same treatment. subprocess_backend gains a background reaper (lazy daemon thread, started on first spawn) that shuts down sidecars idle past OMNIVOICE_SIDECAR_IDLE_TIMEOUT_S (default 300 s; <= 0 disables). The next request transparently respawns one via the existing dead-process relaunch. Safety: the reaper only acts while holding the per-backend lock acquired NON-blockingly, so it can never run mid-op — if an op holds the lock it skips that backend this round. Reuses the idempotent shutdown() (which doesn't take the lock, so no re-entrancy). Each backend tracks last-use and registers in a weak live-set. Scope: subprocess engines only (the heavy, VRAM-holding, process-isolated class). In-process non-default engines and cross-engine VRAM preemption remain TODO — get_active_tts_backend returns a fresh instance per call, so those need an instance-tracking refactor. 6 reaper tests via the stdlib echo sidecar (no torch): kills idle, respawns, skips busy (lock held), recent-use kept, disabled at <=0, ignores dead. The 3 subprocess suites pass together (24). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e8705a106d |
feat(dictation): opt-in NLMS AEC for dictate-over-playback (Wave 8b) (#399)
* feat(dictation): opt-in NLMS AEC for dictate-over-playback (Wave 8b) Dictating while OmniVoice plays audio (TTS preview, dub, video) leaks the loudspeaker signal into the mic, and the streaming ASR transcribes that bleed. Browser echoCancellation varies per platform/webview — it can't be a cross-platform default — so this adds a server-side canceller that behaves identically everywhere. services/aec.py ports Patter's NlmsEchoCanceller (MIT): a time-domain NLMS adaptive filter with a Geigel double-talk detector, warm-up step ramp, and far-end staleness pass-through. /ws/transcribe gains an opt-in '?aec=1[&sr=]' mode: frames are raw int16 mono PCM tagged with a 1-byte prefix (0x00 mic, 0x01 playback reference); the mic is cleaned against the reference before buffering, and the cleaned PCM is muxed via stdlib wave (not ffmpeg). Without the param the protocol and behaviour are byte-for-byte unchanged. Backend ships dark (no new deps — numpy already pinned); frontend far-end streaming is a follow-up. Tests cover echo attenuation, double-talk preservation, cold/stale pass-through, param validation, and the framing helpers — all pure-numpy/stdlib so they skip the torch ASR stack. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(capture_ws): stubs accept the new pcm_sr kwarg _transcribe_buffer/_transcribe_buffer_full gained an optional pcm_sr kwarg for the AEC PCM path; the protocol-test stubs had fixed signatures and raised TypeError on it, so the handler sent 'error' instead of 'final'. Accept **kw in the stubs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e862f0faf0 |
feat(asr): crash-isolated faster-whisper subprocess backend (Wave 4.2) (#393)
* feat(asr): crash-isolated faster-whisper subprocess backend (Wave 4.2) Native ASR engines (faster-whisper / CTranslate2) can segfault on GPU teardown — a process-level crash that kills the whole backend. Running the engine in a child process turns that into a failed job: the sidecar dies, the parent raises a decorated error (engine id + device), and the next request respawns a fresh sidecar. - services/subprocess_asr.py: SubprocessASRBackend reuses SubprocessBackend's wire protocol + lifecycle — including respawn-on-dead-process (_spawn relaunches when the child isn't alive) and GPU-slot acquire/release — adding a 'transcribe' op (the TTS 'generate' surface is stubbed). IsolatedFasterWhisperBackend wraps faster-whisper using the PARENT venv (already a dep — only the process boundary is new); opt-in via OMNIVOICE_ASR_BACKEND=faster-whisper-isolated. - engines/_asr_sidecar/main.py: the faster-whisper runner (stdlib wire protocol; torch/CT2 import lazily so the ready handshake fits the timeout). - engines/_echo/main.py: a 'transcribe' echo op so the round-trip + crash recovery are testable without a real engine. - asr_backend._REGISTRY is now a lazy dict (mirrors the TTS registry) so the isolated backend lists/resolves without importing the subprocess stack unless selected. Tests (echo sidecar, stdlib-only): round-trip, single long-lived sidecar across calls, crash-mid-transcribe → decorated error + backend healthy + next call respawns, registry exposure, generate-not-supported. Spec 7 / parity program Wave 4.2. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(asr): deterministic crash test + drift marker for lazy ASR registry (Wave 4.2 CI) CI surfaced two issues: - The echo crash test relied on the crash-AFTER-reply hook, whose reply may still reach the parent (timing-dependent) — and a leaked OMNIVOICE_ECHO_CRASH from a sibling subprocess test poisoned the non-crash tests. Fix: a deterministic OMNIVOICE_ECHO_CRASH_NO_REPLY hook that exits BEFORE replying (guaranteed dead pipe → decorated error), and the asr fixture clears both crash envs so the round-trip/two-call tests can't inherit a leak. - check-docs-drift's _ASR_MARKER didn't match the new lazy registry line (_LazyASRRegistry({); updated the marker + the self-test fixture. Verified the no-reply crash hook by driving the sidecar directly (reply=None, exit 1); drift self-test + real-repo check green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(asr): allowlist the 'segments' op so transcribe replies aren't dropped (Wave 4.2 CI) The parent's PARENT_INBOUND_OPS frozenset gated inbound sidecar frames but never included 'segments' — the ASR transcribe reply op. _recv() dropped the frame as disallowed, tail-recursed, hit EOF, and returned None, so every transcribe surfaced as a bogus 'sidecar crashed mid-transcription'. TTS ('audio') was allowlisted; ASR ('segments') was missed. Add it (and list 'transcribe' in the informational SIDECAR_INBOUND_OPS), update the exact-shape allowlist test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b1ffdf2387 |
fix(mlx): harden import guards against PyInstaller dylib failures (Wave 4.4) (#390)
MLXWhisperBackend / MLXAudioBackend is_available() caught only ImportError. In a PyInstaller bundle mlx's native dylib/metallib can fail to load even when the package imports, raising OSError/RuntimeError — which would propagate and crash the registry scan instead of reporting the backend unavailable. Broaden to (ImportError, OSError, RuntimeError) so the picker falls back cleanly. 6 tests across all three exception types. The capture ASR path already prefers MLX Turbo on Apple Silicon (get_capture_asr_backend), so this hardening is the remaining slice of Spec 6 / Wave 4.4. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6140f888e1 |
fix(dub+win): dialect↔cinematic guidance loop + WinError 193 ffmpeg validation (#377)
* fix(dub): break the dialect↔cinematic guidance loop (#372, #373) - Cinematic toggle refuses the pick when no LLM endpoint is configured, pointing at Settings → Credentials → LLM endpoint - backend Fast fallback now syncs the quality toggle to 'fast' - the dialect warning no longer fires alongside the cinematic-no-LLM warning (the pair formed the loop), and both messages point at the LLM endpoint settings instead of each other Fixes #372 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ffmpeg): validate resolved ffmpeg/ffprobe actually runs — fall through on WinError 193 (#360, #361, #362) A corrupt or wrong-arch imageio-ffmpeg download (and WindowsApps alias stubs) passes os.path.isfile/shutil.which but explodes at spawn with '[WinError 193] %1 is not a valid Win32 application', killing transcription with an opaque 500. Every resolution step now probes the candidate with '-version' (cached per process), logs the rejected basename, and falls through to the next source. Fixes #362 Fixes #361 Fixes #360 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
10806fea4f |
feat(dictation): optional local-LLM refinement of finals (Wave 2.1) (#363)
Phase 2 of Spec 3, on top of Wave 1.1's deterministic collapse. Prompt design ported from voicebox (MIT): 'text filter, not an assistant' base instruction + three toggleable sections (smart_cleanup, self_correction, preserve_technical) + 7 few-shot examples passed as STRUCTURED chat turns (small local models echo inline examples). Runs through the user's own Ollama/LM Studio/OpenAI-compat endpoint via llm_backend — new additive chat_messages() on the adapter; chat() now delegates to it. Pass-through is the contract: with no LLM configured (backend 'off'), on any error/timeout, or on an empty reply, the raw transcript stands — identical default behavior on every platform. Refinement runs off-thread on FINALS only; the WS final dict gains optional refined_text and the dictation pill pastes refined_text ?? text (raw kept in history). Settings: GET/PUT /api/settings/dictation-refinement (loopback-gated, persisted in the settings table) + a Capture-tab panel with the master switch + per-flag toggles and a 'no LLM configured' hint. 15 new unit tests: prompt sections per flag, structured few-shot message shape, and the full maybe_refine pass-through matrix (off backend, disabled config, LLM failure, empty reply, empty input). Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
93723c2789 |
feat(dictation): collapse Whisper hallucination loops in final transcripts (Wave 1.1) (#356)
Deterministic pre-pass ported from voicebox (MIT, attribution header): word-level (token repeated >=6x, punctuation-normalized) + character-level (2-60-char unit repeated >=6x, catches multi-word and no-space-script loops). Rhetorical repeats below 6 survive; no LLM involved; identical on every platform. Applied to the FINAL text in /ws/transcribe and POST /transcribe — segments keep raw recognition so timings stay truthful. Phase 1 of Spec 3 (docs/competitive-analysis.md); the optional local-LLM refinement pass (phase 2) lands with parity program Wave 2.1 in the same module. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b34dcd9e11 |
Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + integration (#100)
* Phase 4 Plan 04-01: SPIKE-01 GGUF — GO + Wave 1 integration Integrates Serveurperso/OmniVoice-GGUF as a hardware-adaptive default voice-cloning engine, with overridable fallback to the in-process OmniVoiceBackend. Spike confirmed GO: the model is a clean quantization of k2-fsa/OmniVoice (Apache-2.0 + MIT runtime, `omnivoice-lm` custom architecture so it does NOT load in vanilla llama.cpp). Pinned SHAs: * Serveurperso/OmniVoice-GGUF revision: 361609388ae572a820d085185bbbe2a2aac4b30e * ServeurpersoCom/omnivoice.cpp master: 886fc079838ca7400cb2b42b36e2a65aa1daabe8 Implements GGUF-01 (hardware probe) through GGUF-05 (default-engine resolver with graceful fallback). The four `bin/omnivoice-tts-*` artifacts are committed as zero-byte placeholders; the new CI matrix job builds the real binaries per platform from the pinned commit SHA and appends a SHA-256 manifest used by `is_available()` for tampering detection (T-04-01). The macos-14 (Apple Silicon) slot is marked `continue-on-error: true` because omnivoice.cpp publishes no `buildmetal.sh` (Pitfall 1 / Assumption A1) — failure feeds into Task 3's GO/NO-GO call. Quant override is allow-listed against quant_map.json entries only (T-04-05). Argv is composed from typed Path objects rooted in HF_HUB_CACHE; never uses `shell=True`. HF token redaction applies to captured stderr before logging (AUTH-05 / T-04-04). Tests: 36 new (8 hardware-probe + 13 GGUF engine + 6 settings_store quant override + grep gate); 428 passed in full suite vs 402+ baseline. ADR Status stays "Proposed (research-supported)" — Task 3 (human checkpoint) flips to Accepted after CI produces real binaries and a reviewer signs off on the GGUF-06 cross-hardware smoke. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci: install libopenblas-dev on linux-x86_64 omnivoice-tts build The pinned omnivoice.cpp commit (886fc079...) ships a `buildcpu.sh` that passes `-DGGML_BLAS=ON`. ubuntu-latest has no BLAS implementation preinstalled, so the cmake configure step fails with `Could NOT find BLAS (missing: BLAS_LIBRARIES)` and the job exits in 13 s before producing the linux-x86_64 binary. macOS (Accelerate, built in) and Windows (BLAS off by default in the ggml CMakeLists for non-APPLE platforms — the build script doesn't invoke buildcpu.sh on those slots) are unaffected and stay green. Adds a Linux-gated apt step to install libopenblas-dev + pkg-config before the build, restoring cross-platform parity per the CLAUDE.md "default features must work on every platform" rule. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(gguf): constrain ref_audio to project roots — block /etc/shadow on Linux The GGUF engine's `_build_argv` previously validated ref_audio only via `ref_path.is_file()` — i.e. "does this path exist?" That check is platform-dependent: `/etc/shadow` doesn't exist on macOS (rejected naturally), but it IS a real system file on Linux, so the validation silently accepted it. CI's ubuntu-22.04 runner exposed the gap via `test_generate_blocks_freeform_ref_audio`, which exists precisely to guard the "freeform ref_audio path" attack surface. Fix: confine ref_audio to one of three allowed roots before existence checks: - VOICES_DIR (user-saved voice profiles) - DUB_DIR (per-job auto-clones extracted from source video) - tempfile.gettempdir() (browser-upload temp files; existing `cleanup_ref` flow in generation.py) Anything outside those roots → FileNotFoundError, matching the existing failure-mode contract callers handle. Existence check still runs after, so the test's mocked subprocess.run is never reached and the test passes deterministically on all three platforms. Cross-platform parity (per CLAUDE.md 2026-05-20 rule): identical behaviour on macOS / Windows / Linux — the allow-list is computed from core.config which uses platform-specific path resolution but yields the same logical "project tree" on every OS. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * ci(gguf): mark darwin-x86_64 binary build as experimental GitHub's macos-13 (Intel) runner pool is heavily contended — PR #100 queued for 30+ minutes waiting on darwin-x86_64 while every other platform finished in ~1m. Intel Macs are also fading hardware (Apple's platform momentum is entirely on Apple Silicon), and the GGUF engine's runtime already handles a missing binary gracefully (`is_available()` returns False on Intel Mac with a "binary not bundled for this platform" message, same path used for first-launch before any binaries build). `experimental: true` mirrors what darwin-arm64 (Metal) already has — slot still runs and uploads its binary when successful, but a failure or runner backlog no longer blocks merges. Keeps the GGUF engine shippable across the dominant arm64 / Linux / Windows surface without holding the inbox on a slow-runner queue. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
84fffa5409 |
Phase 2 Plan 02-04: Engine Compatibility Matrix API + UI (#99)
* Phase 2 Plan 02-04: GET /engines/{id}/health + gpu_compat + HF mask
ENGINE-06 backend half. Adds the data + spawn-on-demand endpoint the new
Engine Compatibility Matrix UI will consume:
* `gpu_compat: tuple[str, ...]` class attribute on `TTSBackend`, overridden
per backend with reasonable defaults (cuda+mps+cpu for OmniVoice/VoxCPM2;
cpu-only for KittenTTS; mps+cpu for MLX-Audio; etc.). `list_backends()`
serializes it as a list.
* `_HF_TOKEN_MASK_RE` (`hf_[A-Za-z0-9]{30,}`) scrubs the `reason` and
`last_error` fields before they leave the registry — Phase 1's
HFTokenRedactor logging filter does not run on FastAPI response bodies,
so this closes T-02-12.
* `GET /engines/{engine_id}/health` — loopback-gated route that resolves
the backend across tts/asr/llm registries, then either calls
`SubprocessBackend.health_check()` (spawn-and-ping) for subprocess
engines or falls back to `is_available()` for in-process engines.
Returns `{ id, ok, message, latency_ms }`. Engine instances are cached
per-class so repeated checks don't leak atexit hooks or spawn extra
sidecars. The masked-redactor is reapplied on the way out.
Test coverage (tests/backend/api/test_engines_route_shape.py, 11 tests):
* Response shape includes the new fields for every TTS entry
* IndexTTS2 isolation_mode == "subprocess", OmniVoice == "in-process"
* Health route round-trips with mocked SubprocessBackend success
* Health route falls back to is_available for in-process backends
* Unknown engine id → 404
* Non-loopback origin → 403
* Engine instance cache reuses the singleton across calls
* HF tokens leaked into is_available() / health_check() are masked
in both the /engines and /engines/{id}/health response bodies
Existing tts_backend_registry shape test updated to include `gpu_compat`.
Full suite: 402 passed, 0 failures (up from 391+ baseline).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Phase 2 Plan 02-04: EngineCompatibilityMatrix UI + Settings wiring
ENGINE-06 frontend half. Mounts a new component on Settings → Engines
that surfaces, end-to-end, the data shape Plan 02-01 + Plan 02-03 added
to the backend registry:
* `frontend/src/components/EngineCompatibilityMatrix.jsx` (270 lines) —
semantic <table> with role=row/cell so RTL queries work; one row per
registered backend. Columns:
- Engine name + install hint + Last error line
- Install state badge (Available / Unavailable + inline reason)
- GPU compat chips (CUDA / MPS / ROCm / CPU with colored variants)
- Isolation mode badge (subprocess for IndexTTS, in-process for the
rest — makes the Phase 2 architectural shift legible to users)
- "Test engine" button → `/engines/{id}/health` round-trip; renders
latency in ms inline next to the button; disabled while inflight;
5 s cooldown to prevent click-storms.
Mount does NOT auto-test any engine — per the plan's Open Question #2,
spawning sidecars is gated on user action.
* `frontend/src/components/EngineCompatibilityMatrix.css` — minimal
styling that reuses chrome tokens; chip colors per GPU target.
* `frontend/src/api/engines.ts` — `getEngineHealth(id)` client function
wraps the new backend route through the shared apiJson helper.
* `frontend/src/api/types.ts` — extends EngineBackend with optional
`isolation_mode`, `last_error`, `install_hint`, `gpu_compat` so the
TypeScript surface tracks the backend wire shape, and adds
EngineHealthResponse.
* `frontend/src/pages/Settings.jsx` — replaces the hand-rolled Engines
table inside EnginesTab with `<EngineCompatibilityMatrix family="tts"
onSelect={...} />`. selectEngine still wires up the picker; the
matrix's onSelect prop renders the Use button per row when provided.
Removes the now-unused FAMILY_META local map.
Test coverage (`frontend/src/test/EngineCompatibilityMatrix.test.jsx`,
8 tests via vitest):
* Renders one row per backend with documented columns
* isolation_mode badge: subprocess for IndexTTS2, in-process for
OmniVoice / KittenTTS
* GPU compat chips: omnivoice → cuda/mps/cpu; kittentts → cpu only
* Unavailable rows render the failure reason inline
* last_error line renders below status when populated; masked HF
token sentinel survives verbatim
* Test engine click fires getEngineHealth(id) and renders latency_ms
* Test button disabled while inflight; second click is a no-op
* Failure path (ok=false) renders a failure marker
Frontend suite: 65 passed (8 new). Lint: 0 new errors. typecheck:ci: clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Phase 2 Plan 02-04: SUMMARY
Recap of Engine Compatibility Matrix delivery — backend route +
gpu_compat metadata + HF-token redaction, frontend EngineCompatibility-
Matrix component, full test counts, deviations, gpu_compat confidence
matrix, frontend test-runner command notes for Phase 6 CI.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
c3695e1668 |
Phase 2 Plan 02-03: IndexTTS on SubprocessBackend (closes #42) (#98)
Migrates IndexTTS-2 off the in-process import path and onto the SubprocessBackend primitive shipped in Plan 02-01. Closes issue #42 with a structural fix — the parent's transformers>=5.3 and IndexTTS's transformers<5 now live in separate OS processes and can never collide. * New: backend/engines/indextts/ — sidecar package (__init__.py hosts IndexTTS2Backend, main.py is the sidecar entrypoint, bootstrap.py owns the 3-step venv probe + lazy uv-based bootstrap). * services.tts_backend: IndexTTS2Backend's in-process body removed; registry resolves the class lazily via a _LazyRegistry indirection + PEP 562 __getattr__ re-export. This breaks the import cycle that arose when both subprocess_backend and tts_backend tried to import each other at module load. * docs/engines/indextts.md: install walkthrough + venv resolution order + common errors (linked from is_available()'s unavailable message). * tests: - test_indextts_backward_compat.py (8) — probe priority, no-spawn discipline, HF cache marker preservation (ENGINE-07). - test_indextts_sidecar.py (17) — subclass shape, isolation_mode, parent-side emotion arbitration (vector/audio/text/description), coexist-with-OmniVoice (headline #42 closure), env forwarding. - tests/fixtures/mock_indextts_sidecar.py — stdlib-only sidecar mimicking the production wire protocol; emits 1 s sine wave. - test_issue_fixes.py: two obsolete in-process-conflict tests rewritten to assert the new subprocess contract (no indextts.* import in the parent). Hard constraints honored: backend/services/sonitranslate.py and gpu_sandbox.py are untouched (D1 / D4). Existing v0.2.7 users with OMNIVOICE_INDEXTTS_DIR and a populated HF cache reach a working generation with zero re-download and zero re-install. 44 tests pass across the four exercised files. Full suite: 391 passed, 10 skipped, 13 xfailed, 1 xpassed in 57 s. Smoke: 4 passed. Closes #42. Requirements: ENGINE-02, ENGINE-03, ENGINE-04, ENGINE-07. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
0fc5ea6cf3 |
Phase 2 Plan 02-01: SubprocessBackend primitive (Wave 1 of Phase 2) (#97)
* Phase 2 Plan 02-01: SubprocessBackend primitive + echo sidecar + ENGINE-05 wrap
Lands the durable SubprocessBackend primitive — the architectural keystone
that Plans 02-03 (IndexTTS migration), Phase 3 (Supertonic-3), and
Phase 4 (GGUF / Singing) plug into.
Files added:
- backend/services/subprocess_backend.py — base class owning spawn,
shutdown, _send/_recv (length-prefixed JSON), GPU-slot acquire-release,
atexit teardown, stderr drain, op allowlist (T-02-04), and 64 MB
frame cap (T-02-01). No multiprocessing — subprocess.Popen
exclusively so subclasses can target a *different* venv's interpreter
(Locked Decision D4 / Pitfall 1).
- backend/engines/_echo/main.py — permanent CI regression sidecar.
Stdlib-only, runs under the parent's sys.executable. Implements
ready/ping-pong/synthesize/shutdown plus test-only probe_env and
emit_unknown ops for env-forwarding and op-allowlist tests. DO NOT
DELETE — the round-trip test depends on this file.
- tests/backend/services/test_subprocess_backend.py — 13 tests:
round-trip, health_check, no-zombie, shutdown idempotency, env
forwarding (HF_TOKEN/HF_HOME/HF_ENDPOINT/HF_HUB_CACHE), oversize
frame, short read, op-allowlist drop, op-allowlist constant shape,
sidecar-crash recovery, no-multiprocessing grep gate, MAX_FRAME_BYTES.
- tests/backend/services/test_tts_backend_registry.py — 6 tests for
list_backends() resilience + shape + isolation_mode + last_error
caching + existing-engines preservation + install_hint passthrough.
Files modified:
- backend/services/tts_backend.py:
* Adds module-level _LAST_ERRORS dict for ENGINE-06.
* Rewrites list_backends() to wrap each is_available() in try/except
so one broken engine cannot blank the picker (ENGINE-05).
* Adds last_error + isolation_mode keys to each response entry
(ENGINE-06 UI in Plan 02-04 consumes via the same /engines route).
* Uses a duck-typed _is_subprocess_isolated marker rather than
issubclass(cls, SubprocessBackend) because test fixtures (token
resolver suite) purge sys.modules["services"] between tests and the
re-imported SubprocessBackend would be a different class object.
Threat-model mitigations (Plan 02-01 frontmatter):
T-02-01 DoS via length-prefix → MAX_FRAME_BYTES = 64 * 1024 * 1024
T-02-02 GPU slot leak on sidecar death → try/finally in generate
T-02-03 token bytes in stderr → drained via parent logger
(HFTokenRedactor from Phase 1 already on root)
T-02-04 unknown ops from compromised sidecar → PARENT_INBOUND_OPS
allowlist, unknown frames logged and dropped
T-02-05 Tauri group-kill scope → start_new_session=True on Unix /
CREATE_NEW_PROCESS_GROUP on Windows
Verification:
- 337 passed, 6 skipped, 12 xfailed, 1 xpassed (full suite,
`uv run pytest tests/ --ignore=tests/manual`)
- All 19 new tests pass on macOS Apple Silicon
- Smoke tests still pass: `uv run pytest tests/smoke/ -q` → 4 passed
- SoniTranslate untouched (D1 locked decision)
- Zero new Python dependencies
Closes part of ENGINE-01 + ENGINE-05.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(02-01): plan summary — public API, invariants, deviations
Documents the SubprocessBackend public API so Plan 02-03 (IndexTTS) and
Phase 3 (Supertonic-3) authors don't need to re-read the source.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
||
|
|
c6e9bbc191 |
Phase 2 Plan 02-02: audio I/O hardening + WAV-export correctness (#96)
* Phase 2 02-02: add _safe_torchaudio_save + _safe_soundfile_write helpers
Centralizes WAV/audio writes through a single audited path that defends
against the four documented torchaudio.save failure modes (CUDA/MPS
tensor, non-contiguous, out-of-range, wrong dtype) AND the torchaudio
2.9+ TorchCodec-delegation behavior drift.
* services/audio_io.py:_safe_torchaudio_save now performs:
- .cpu() move (torchaudio cannot serialize CUDA/MPS)
- dtype coercion to torch.float32
- .clamp(-1.0, 1.0) (out-of-range = silent clipping on some backends)
- .unsqueeze(0) for 1D (mono) inputs
- .contiguous() (torch.cat of slices = non-contig = silent corruption)
- explicit encoding="PCM_S/PCM_F" + bits_per_sample so future
torchaudio backend selection cannot drift the on-disk format
- format passthrough for wav/flac/mp3/ogg with encoding-kwarg fallback
for older codec builds
* services/audio_io.py:_safe_soundfile_write — sibling helper for the
one sf.write call site (dub_core.py). Applies the same dtype/contig/
range checks before delegating to soundfile.write.
* services/audio_io.py:atomic_save_wav (existing P0 helper) now
delegates the actual encode to _safe_torchaudio_save so atomicity
and correctness compose: every byte that lands at the target path
was produced by the audited helper.
* tests/backend/services/test_audio_io.py — 29 tests (25 pass + 4
skipped for MPS dtype incompatibility): parametric round-trip across
dtype x device x contiguity, plus out-of-range clamp, format
passthrough, in-memory buffer, empty-tensor rejection, 1D auto-
unsqueeze, and a smoke check that atomic_save_wav inherits the
safety guarantees.
No new Python dependencies. SoniTranslate untouched (D1 locked).
Refs BUG-01 / #48.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Phase 2 02-02: migrate router audio writes through audited helpers
Migrates all 12 grep-audit bare audio-write call sites in
backend/api/routers/ to route through services.audio_io. Closes the
last surface area of BUG-01 / #48 that the P0 atomic-write commit
(
|
||
|
|
c32041289d |
Phase 1 Wave 3: AppImage launcher + .deb ffprobe + Docker LAN + Gatekeeper probe (closes #54, #56, #76, #80) (#93)
* fix(appimage): conditional WEBKIT_DISABLE_COMPOSITING_MODE launcher (#56) WebKitGTK 2.44.x and 2.46.x have a compositing-path regression on Wayland that blanks the AppImage's first paint on Fedora 44 / Ubuntu 24.04. Setting WEBKIT_DISABLE_COMPOSITING_MODE=1 forces the software fallback that works, but blindly setting it on healthy WebKit versions (2.48+) regresses those. This wave adds a conditional AppRun launcher that detects the WebKit version via pkg-config and only sets the env var on the broken ranges (plus a fail-safe when pkg-config is absent or the version is unknown). The launcher is injected into Tauri's AppImage staging dir via a beforeBundleCommand hook — see .planning/decisions/apprun-strategy.md for the spike outcome and rationale (Strategy B chosen). Phase 1 Wave 3 — Plan 01-03 Task 1. Closes #56 frontend half. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(deb): relocate bundled ffprobe out of /usr/bin to avoid conflicts (#76) Prior versions placed the bundled ffprobe at /usr/bin/ffprobe via Tauri's externalBin, which overwrites the system ffprobe on Ubuntu 26.04 and collides with apt-installed media-package ffprobe. Relocate the .deb-bundled ffprobe to /usr/lib/omnivoice-studio/bin/ffprobe via bundle.linux.deb.files, plus defensive maintainer scripts: - preinst: ensure target dir exists for upgrade flows - postinst: remove legacy /usr/bin/ffprobe ONLY when dpkg confirms our package owns it (never touches a user's distro ffprobe) - postrm: clean up the relocated path tree on purge/remove Rust side (tools.rs::resolve_ffprobe) now probes the new path on Linux, and backend spawn (backend.rs) carries both FFPROBE_PATH (legacy alias) and OMNIVOICE_FFPROBE_PATH (canonical) into the backend env. Python side (ffmpeg_utils.resolve_ffprobe) reads OMNIVOICE_FFPROBE_PATH first, falls back to FFPROBE_PATH, then to shutil.which("ffprobe"). 6 new unit tests cover the env-cascade resolution. Phase 1 Wave 3 — Plan 01-03 Task 2. Closes #76. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(frontend): centralised apiBase resolver for Docker LAN access (#80) Docker / LAN browser users hit the preview API at the LAN host's IP, not their local machine — the prior frontend/src/utils/media.js:20 hardcoded http://localhost:3900, which from a LAN client resolved to the client machine itself. Centralise via frontend/src/utils/apiBase.ts: 1. VITE_OMNIVOICE_API override (Docker compose / dev) always wins. 2. Tauri webview → http://localhost:3900 (unchanged behaviour). 3. Plain browser → ${window.location.protocol}//${window.location.hostname}:3900 (follows the page's origin — closes #80). 4. SSR / no-window → http://localhost:3900 (safe fallback). Grep-sweep confirmed media.js:20 was the only hardcode site (Assumption A4 in 01-RESEARCH.md verified). 6 new vitest cases cover the resolver. Phase 1 Wave 3 — Plan 01-03 Task 3. Closes #80 frontend half. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(backend): macOS Gatekeeper quarantine probe + INST-01 guard (#54) Adds backend/core/gatekeeper_detect.py which walks up from sys.executable to find the .app bundle and runs `xattr -l` to check for the quarantine extended attribute (com.apple.quarantine). On detection, the lifespan startup probe logs a structured warning and emits a system_error event through the existing event bus with error_class="GATEKEEPER_QUARANTINE", which Wave 2's React ErrorBoundary turns into a docs deeplink. Detection is informational only — we never auto-run `xattr -cr` (the app itself is quarantined and cannot fix its own state per Anti-Pattern in 01-RESEARCH.md). Users get a clear pointer to the workaround docs. GET /system/quarantine-status exposes the structured payload so the frontend can poll on first load. INST-01 (setuptools>=75.0 pin from PR #62) gains a PR-time guard in tests/backend/test_pyproject.py + a user-observable smoke check in scripts/smoke-test.sh (pkg_resources + whisperx import). 7 gatekeeper tests + 1 pyproject test added — all pass. Phase 1 Wave 3 — Plan 01-03 Task 4. Closes #54 backend half (Wave 2 owns the docs page + ErrorBoundary deeplink wiring). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
4a6b978df9 |
Phase 1 Wave 1: HF token persistence + redactor (closes #35) (#91)
* feat(01-01): encrypted settings store + alembic migration (AUTH-02, T-01-01) Adds the SQLite-backed encrypted settings store that Phase 1 token resolver will read from. Closes the at-rest plaintext risk for HF tokens (T-01-01). - backend/services/settings_store.py: get_hf_token / set_hf_token / clear_hf_token using Fernet symmetric AEAD. Stored value column never contains the literal "hf_" substring. - backend/services/_secret_key.py: per-install Fernet key derived via scrypt(machine-id + 16-byte random salt). machine-id resolution covers macOS (ioreg IOPlatformUUID), Linux (/etc/machine-id and dbus fallback), Windows (HKLM Cryptography MachineGuid via winreg). Final fallback to hostname+user with a warn log. - backend/migrations/versions/0001_phase1_settings_table.py: alembic migration adding `settings(key, value, updated_at)`. Idempotent — checks for an existing table so fresh installs (where _BASE_SCHEMA already created it) and v0.2.7 upgrades both succeed. - backend/core/db.py: _BASE_SCHEMA grows the settings table for fresh installs; init_db() now runs `alembic upgrade head` after the CREATE. - backend/migrations/env.py: honours an externally-set sqlalchemy.url so tests can point alembic at a fixture DB; falls back to core.config DB_PATH for production. - pyproject.toml: cryptography>=41 added explicitly (RESEARCH.md Assumption A1 was checked at execute-time and proved false; the dep was not present transitively, so the install would fail without this). Tests (10 cases, all green): - Round-trip encryption + plaintext-leakage check (T-01-01 invariant) - Salt persistence across clear/set cycles - InvalidToken decrypt path returns None (Open Question #5 resolution) - Concurrent reads consistent under sqlite WAL - Alembic upgrade on a hand-built v0.2.7 fixture DB preserves all existing tables + seeded rows (CLAUDE.md backward-compat constraint) - Alembic downgrade -1 drops only the settings table Refs #35. * feat(01-01): 3-source HF token resolver + log redactor + 5 read sites patched Closes the #35 bug class (bare os.environ.get('HF_TOKEN') reads) by routing every backend HF-token consumer through one resolver, and mitigates T-01-02 (info disclosure via logs) by stripping `hf_[A-Za-z0-9]{30,}` substrings from every log record at the root logger. backend/services/token_resolver.py: - resolve(skip) — 3-source cascade (App → Env → HF-CLI), each source validated via huggingface_hub.whoami(); first valid wins. - on_401(active) — invalidate cache and re-resolve skipping the source that just 401'd (AUTH-06). - state() — three SourceState rows for the Settings UI: set, masked preview (hf_…<last 3>), whoami_user, whoami_ok. - save_app_token / clear_app_token — wraps settings_store + calls huggingface_hub.login(add_to_git_credential=False) per Pitfall #2. - 300-second whoami cache so repeated Settings-page renders don't hit the HF API. backend/core/logging_filter.py: - HFTokenRedactor(logging.Filter) — regex `hf_[A-Za-z0-9]{30,}` so real tokens are masked but `hf_hub` / `hf_token` literals survive. - install_redaction_filter() — idempotent attach to root + every handler. backend/main.py: install the redactor at startup, BEFORE the file handler is added. Re-installed after the file handler attaches so the handler-attached filter list includes it too. Read-side call sites patched (per Pitfall #1 — every HF token read must flow through token_resolver.resolve()): - backend/api/routers/dub_core.py:540 (the original #35 site) - backend/api/routers/system.py:38 (_has_hf_token notification) - backend/services/model_manager.py:480 (diarization pipeline auth) - backend/services/sonitranslate.py:143 (Popen env for SoniTranslate child) - backend/services/sonitranslate.py:217 (gradio_client predict call) New endpoint: - GET /system/hf-token/state — returns the 3-source cascade state with masked tokens for the Wave 2 Settings UI panel. Grep gate confirmed clean: zero `os.environ.get("HF_TOKEN")` reads remain outside token_resolver.py. Tests (17 new cases, all green): - tests/backend/services/test_token_resolver.py: priority cascade, 401 skip mid-resolve, on_401 fallback, state() shape, save+login invariant (add_to_git_credential=False), HUGGING_FACE_HUB_TOKEN alias acceptance. - tests/backend/core/test_logging_filter.py: msg + args redaction, multi-token redaction, non-string args pass-through, short-token literals preserved, install_redaction_filter idempotence. Refs #35. * feat(01-01): Settings hf-token API endpoints + subprocess env injection (AUTH-03/04) Backend half of the Wave 2 Settings → API Keys UI plus the AUTH-04 subprocess env-injection invariant. backend/api/routers/settings.py: - POST /api/settings/hf-token — body {token: str} → save_app_token - DELETE /api/settings/hf-token — also_clear_hf_cli query → clear_app_token - GET /api/settings/hf-token/state — same shape as token_resolver.state() All three are gated by `Depends(require_loopback)` at the router level (threat T-01-03 mitigation; non-loopback Host → 403). backend/main.py: router mounted alongside existing API routers. Subprocess env injection (AUTH-04, threat T-01-04 disposition=accept): - backend/services/sonitranslate.py already updated in Task 2 to read via token_resolver.resolve() and inject HF_TOKEN + YOUR_HF_TOKEN into the SoniTranslate child env block. - backend/services/gpu_sandbox.py: NOT patched — the GPU sandbox runs in-process TTS generation that uses the parent's already-loaded HF state. Adding env injection there is a no-op (parent and child share state via multiprocessing.Pipe before any HF API call). - backend/services/model_manager.py:480 (Task 2): resolves in-process, no subprocess crosses here. - backend/api/routers/exports.py: subprocess.Popen calls only spawn `open` / `explorer` / `xdg-open` — file-manager launchers with no HF needs. Skipped per Task 3 conservative-patching rule. So the canonical AUTH-04 site for this milestone is sonitranslate.py. Future SubprocessBackend work in Phase 2 will inherit the same pattern. Tests (8 new cases, all green): - tests/backend/test_engine_spawn_token.py * POST /hf-token loopback → 200 + state.active == "app" * POST /hf-token non-loopback → 403 ("loopback origin required") * DELETE /hf-token clears settings_store + state.active == None * GET /hf-token/state returns 3 source rows in priority order * GET /hf-token/state non-loopback → 403 * env block contains HF_TOKEN + YOUR_HF_TOKEN when resolver returns one * env block does NOT contain an injected empty HF_TOKEN when resolver returns None * source-level check that backend/services/sonitranslate.py still reads via token_resolver.resolve() (regression guard against silent reverts of the AUTH-04 wiring) Full Wave 1 test suite: 35/35 green. Phase 0 smoke tests still green. Refs #35. * docs(01-01): SUMMARY + STATE update for Phase 1 Wave 1 completion Records execution outcome of the 3-task plan: 10 files created, 9 modified, 35 new test cases, 5 read sites patched, grep gate clean. Documents the two Rule-3/Rule-2 deviations applied (cryptography dep, env.py URL override), the subprocess-launcher inventory for Phase 2, and the known stray edit to the main repo's pyproject.toml that needs a one- line user action to revert. Updates STATE.md current-position table, progress bar, and open TODOs to point at Wave 2 (Plan 01-02) and Wave 3 (Plan 01-03) as the next steps. |