fd7d20fe1e53d4e09b33fa59bb47760ba603b520
178
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d14b37fab2 |
fix(dub): the speaker-count hint is honored on every diarization path + clone-purity guard (#952)
* fix(dub): the speaker-count hint is honored on every diarization path + clone-purity guard
The dub "Speakers" count reached _diarize() and then died on 3 of its 4
branches, so setting it changed nothing, speakers blended, and auto-clones
were cut from mixed-speaker audio ("made up" voices):
- FunASR inline-turns shortcut returned before the hint was ever consulted
→ now an explicit num_speakers routes the job through pyannote (the one
engine that honors an exact count); turns stay the fast path only when no
hint is set, and remain the fallback (with an honest "hint ignored"
warning) when pyannote can't load or crashes mid-run.
- pyannote-unavailable fallback used a hardcoded 2-speaker silence-gap
heuristic → assign_speakers_heuristic now takes num_speakers and cycles N
labels on gap boundaries (1 → single speaker; None → legacy alternation),
and the existing diarization warning says the hint is only approximately
honored.
- pyannote-crash fallback dropped the hint the same way → same treatment.
No branch drops the hint silently anymore: every degraded path extends the
existing `warning` SSE payload (detail + a machine-readable speaker_hint
field) that the frontend already renders.
Parity + purity:
- POST /dub/transcribe/{job_id} (the CLI's endpoint) gains the same clamped
num_speakers query param, forwarded to pyannote and the heuristic; the
omnivoice-dub CLI gains --speakers N.
- Clone-purity guard: _pick_reference_slices rejects sub-1.5s slices, prefers
slices not temporally adjacent (<0.3s) to another speaker's turn (scoring
preference, not a hard filter), and extract_speaker_clones skips extraction
entirely when labels came from the heuristic (labels_source kwarg threaded
from _diarize; missing kwarg keeps the old behavior) — with a user-facing
warning pointing at Settings → Models → pyannote.
Tests: fail-before/pass-after coverage in tests/test_speaker_hint.py (all
four _diarize branches driven through the real SSE stream), clone-purity
guards in tests/test_speaker_clone_purity.py, heuristic hint semantics in
tests/test_segmentation.py.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): add the speaker-hint + clone-purity fix under [Unreleased] (#952)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
8bbab3fcc3 |
fix(translate): Dub LLM engine runs on the configured LLM provider (new dub_translation skill) (#944)
* fix(translate): the Dub LLM engine now runs on the configured LLM provider Picking "LLM (OpenAI-compatible)" in the Dub tab read only the raw TRANSLATE_* env vars — completely bypassing Settings → LLM Providers, so a provider the user had configured AND tested in-app silently didn't power the engine (empty key → raw 401 per segment). The Cinematic refiner was already rewired through LLM Skills (#910/#912); this closes the gap for direct LLM translation: * new "dub_translation" LLM skill (Settings → LLM Skills) — per-skill provider override → global active provider, same resolution as every other skill; disabled == unconfigured, no new degradation modes * the provider=openai branch resolves through resolve_skill_client(); TRANSLATE_BASE_URL/TRANSLATE_API_KEY/TRANSLATE_MODEL stay working as the power-user override (env-only setups see zero behavior change, except the stale gpt-3.5-turbo default is now gpt-4o-mini, matching the cinematic path) * per-segment calls are now bounded by the LLM timeout (45s default via OMNIVOICE_LLM_TIMEOUT) instead of the SDK's 600s default * fully unconfigured → an up-front actionable 400 naming Settings → LLM Providers / LLM Skills instead of a per-segment 401 * provider-store keys are resolved into the error scrubber so a provider echoing the key can't leak it (parity with the env-key scrub) * translation_engines registry: honest notes + a configured/configured_via stamp on LLM entries so the Engine dropdown can show ready-vs-needs-setup before the user clicks Translate Tests: 4 new (skills-resolved client wins with its model+timeout; 400s name the right settings page for no_provider vs disabled; env fallback keeps working incl. TRANSLATE_MODEL); skills registry coverage updated; existing openai-branch tests routed deterministically through the env branch via the shared fake helper. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): add the dub-translation provider wiring under [Unreleased] (#944) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ba4f64240a |
fix(engines): classify sherpa "model not set" as a config error, gate the engine on its model dir (#919) (#934)
A user selected the sherpa-onnx TTS engine and got a 500 that read "TTS engine stopped mid-generation. This usually means it ran out of memory. Try the Flush button…" — when the real cause was a pure setup problem: "OMNIVOICE_SHERPA_MODEL not set. Point it to a sherpa-onnx TTS model directory (containing model.onnx + tokens.txt)." Same misclassification class as #880/#893, which tightened the OOM catch-all on the generation path — but the engine-not-configured case still fell through to memory. Two layers, fixing the whole class: 1. Error classification (backend/api/routers/generation.py): a new `_is_config_failure()` recognizes "required engine model path / env var not set" over the whole exception chain (OMNIVOICE_* named with "not set"/"point it to"/"set omnivoice_…", sherpa's "no model.onnx found in", "not configured", "venv not found. set" for the dedicated- venv opt-ins). `_oom_friendly_reraise` checks it BEFORE the OOM branch and re-raises actionable setup guidance that names the variable, points at Settings → Engines, and never mentions memory or Flush. Generalizes to sherpa/Confucius4/dots/MOSS and any future env-gated engine. 2. Engine gating (backend/services/tts_backend.py): SherpaOnnxBackend ships no bundled model, so is_available() now gates on OMNIVOICE_SHERPA_MODEL (set + contains model.onnx) — like the other path-configured opt-in engines — returning False with an actionable reason instead of "ready", so the picker marks it unavailable-with-a- reason rather than selectable-but-broken. Added the copy-paste setup snippet for the Compat Matrix. Backward-compatible: a correctly configured OMNIVOICE_SHERPA_MODEL keeps the engine available. Tests (fail-before/pass-after): config-classification of the sherpa "model not set" error and the wider not-configured class (no "out of memory"/"Flush"); is_available gating on the env var + model.onnx and the setup-snippet registration. Fixes #919 Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5bd8968aea |
feat(engines): real synthesis "Self-test" + copy-paste setup snippet for opt-in engines (#930)
* feat(engines): real synthesis "Self-test" + copy-paste setup snippet for opt-in engines Builds on #905's Engines-settings fixes (verified still green: license dialog mounts, matrix reloads on select, cpu_fallback routing toast, cpu-native → cpu_only). Two enhancements, no #905 behavior touched. Real "Self-test" for in-process TTS engines ------------------------------------------- The existing /engines/{id}/health probe only imports the package and reports "deps OK" for in-process engines — it never proves the engine can emit audio. New POST /engines/{id}/selftest runs a *tiny real synthesis* from a fixed short ASCII phrase and reports ok + duration + sample-rate + sample count, proving the engine actually produces audio. Guardrails keep it cross-platform-identical and CPU-cheap: TTS + available + in-process only, bounded wall-clock timeout (OMNIVOICE_SELFTEST_TIMEOUT_S, default 90s) that returns ok=false/timed_out instead of hanging the panel, a process-wide lock so a click-storm can't stack model loads, loopback-gated, and only ever on user click (never on load). The Compat Matrix gains a "Self-test" button (with cooldown) that renders "0.82s @ 24 kHz in 820 ms". HF tokens in a synth error are redacted like the health route. Verified end-to-end: kittentts synthesized 89,200 samples @ 24 kHz. Copy-paste setup snippet for path-gated opt-in engines ------------------------------------------------------ IndexTTS / MOSS-v1.5 / dots.tts / Confucius4 gate on an OMNIVOICE_*_DIR env var. list_backends() now emits a single-sourced `setup_snippet` (the exact `export VAR=/path/...` line) surfaced with a Copy button inside the matrix's "Why unavailable?" disclosure, so users don't reconstruct it from the docs. Also tightened the incomplete SelectEngineResponse TS type to include the routing echo (routing_status/effective_device/routing_reason) the post-select toast already reads at runtime. Tests: backend selftest success/subprocess-reject/unavailable/unknown/loopback/ exception-capture/timeout/HF-redaction + setup_snippet shape; frontend self-test render, timeout marker, subprocess+ASR gating, setup-snippet render. New route added to the API route snapshot. Full vitest (808) + backend engine/routing/asr/ route-inventory/no-CJK green; lint 0 errors; format + typecheck:ci clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(test): allow setup_snippet key in list_backends shape assertion The engine self-test PR added setup_snippet to each backend entry but only updated the route-shape test; test_list_backends_shape strict-asserts the key set. Add setup_snippet there too. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8f71c90f20 |
feat(settings): LLM Skills — per-feature enable/route control for every LLM call (#912)
New Settings → System → LLM Skills area: every LLM-powered capability
(Cinematic & Autofit translation, speech-rate slot fitting, glossary
auto-extract, direction parsing, dictation cleanup) becomes a "skill" the
user can toggle or route to a specific provider (local Ollama/LM Studio vs
a remote key) instead of everything riding the one global active provider.
Backend:
- services/llm_skills.py — skill registry + settings_store persistence
(llm_skill.<id>.enabled / .provider), resolution precedence
override > active > none, resolve_skill_client() (OpenAI-compat client
bound to the effective provider; None when disabled/unconfigured) and
skill_backend() (OffBackend when disabled — the exact no-LLM object every
caller already degrades on).
- All five consumption points wired through the registry; a disabled skill
degrades exactly like "no LLM configured" today (Fast translation
fallback, refinement pass-through, heuristic direction parse, no-llm slot
fit, 503 on glossary auto-extract). No new degradation modes; defaults
(enabled + no override) keep existing setups byte-identical.
- OpenAICompatBackend gains an optional bound provider (None = active, the
historical behavior).
- GET /api/settings/llm-skills + PUT /api/settings/llm-skills/{skill_id}
(404 unknown skill/provider); route snapshot updated.
Frontend:
- LLMSkillsPanel (Sparkles, next to LLM Providers): one row per skill —
i18n name/description, enable toggle, provider Select ("Use active
provider" + configured providers, local ones tagged), ready /
needs-setup badge linking to LLM Providers. All strings via t()
(settings.llmskills_*).
Tests: 30 backend (precedence, per-consumption-point disabled semantics,
endpoint round-trips, validation) + 4 panel render/PUT tests. Docs:
translation-engines.md gains an LLM Skills section.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
af6690840e |
fix(translate): run Cinematic/Autofit on every engine (incl. default Argos), bound the fit pass, scrub provider errors (#910)
P0 — Cinematic/Autofit silently no-op'd on argos/nllb/openai. Those three branches returned BEFORE _maybe_cinematic, so only the deep_translator fall-through reached the refine/fit pass. A user on the DEFAULT Argos engine who picked Cinematic/Autofit got plain Fast output with a success toast and no quality_used/cinematic_skipped/rate_ratio. All three now route through _maybe_cinematic. provider=openai is already an LLM translation, so it skips the reflect/adapt re-refine (new already_llm flag) but still stamps rate-ratio badges and runs the Autofit fit pass; the dialect it baked into its translate prompt is now reported applied. P1 — the Autofit fit pass ran one blocking adjust_for_slot per segment in the merge loop, OUTSIDE any budget (a 50-seg dub vs a slow provider spun ~50×timeout unbounded). New speech_rate.adjust_for_slot_many fans it out concurrently under a wall-clock deadline SHARED with the cinematic refine; segments still running at the deadline degrade to their literal with rate_error='fit-budget'. Also set max_retries=0 on the OpenAI clients used for translate/refine/fit so a 429 + Retry-After can't sleep through the budget. P2 — glossary auto-extract's no-LLM message now points at Settings → LLM Providers (was the stale TRANSLATE_BASE_URL/TRANSLATE_API_KEY). Provider error bodies on the glossary auto-extract, the OpenAI translate-segment path, and the DeepL/Microsoft translate-segment path are now scrubbed (core.scrub.scrub_provider_error) — they could echo the API key / a user_id. DubTab re-polls LLM availability on window focus / visibility so configuring a provider in Settings lifts the Cinematic gate without a remount. Documented LLM_DEFAULT_PROVIDER in docs/dubbing/translation-engines.md. Tests: fail-before/pass-after for argos+cinematic (refine runs), argos+cinematic no-LLM (cinematic_skipped), argos Fast (rate_ratio stamped), openai+autofit budget bound, and provider-error scrubbing on the translate + glossary paths. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
75864a597f |
fix(dictation): refinement never stalls a final (~51s→≤4s), REST polish parity, real ASR preload reuse (#911)
P0 — Refinement blocked every dictation final with no timeout. With refinement auto:true and a slow/dead LLM endpoint, maybe_refine ran unbounded and blocked the final send in all three capture_ws handlers (~51s measured; the pill hung "Transcribing…" until the widget's 15s fallback fired). Fix the class: a hard, env-tunable budget (OMNIVOICE_REFINE_TIMEOUT_S, default 4s) via a new maybe_refine_async — a slow/dead endpoint now falls back to the unrefined (but polished) text within the budget and can NEVER delay the final beyond it. The LLM HTTP call is bounded to the same budget so the orphaned worker unwinds instead of holding a connection for the client's full 45s. Refinement is now also fully best-effort in the legacy handler (it can't turn a good final into an error frame). P1 — REST /transcribe lacked polish parity. capture.py never applied polish_text, so REST returned raw "…test" while the WS returned "…test." Apply text_polish.polish_text to `text` and `refined_text` (segments stay raw), so the widget POST fallback and MCP/CLI callers match the live socket. P1 — The #888 "instant first dictation" preload was a no-op. The preload called warmup() only `if hasattr`, but SherpaDictationBackend had none, and the WS handlers built a FRESH backend per session so a warm singleton wasn't reused. Add SherpaDictationBackend.warmup() (builds the recognizer) and share one warm recognizer per model id across sessions (get_sherpa_dictation_backend, same invalidation + a shared lock as the capture singleton); each session keeps its own decode stream. First dictation no longer pays the 1.3–2.5s load. P1 — llm_ready is a lie (feeds the P0). It only means "an endpoint is configured", so a placeholder key reads as ready. The P0 timeout makes a dead endpoint harmless; add last_refine_status so RefinementPanel flags a configured-but-failing LLM and links to LLM Providers → Test. Regression tests (fail-before/pass-after): slow-LLM WS final arrives < budget; maybe_refine_async hard timeout + status; REST polish parity + refined_text polish; warmup builds the recognizer and a second session reuses it; the panel honesty note. Backend refinement/capture_ws/capture/sherpa suites, CJK + route inventory gates, full vitest (733), lint (0 errors) and format all green. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
16294fed44 |
feat(updates): data-safe updates — pre-migration DB backups, guarded venv heal, release notes + changelog reader (#909)
Backend: - core/db_backup.py: WAL-safe SQLite snapshot to omnivoice.db.backup-<version>-<n> before pending alembic migrations run; keep newest 3, prune older; skip >500MB with a log line. Restore is never automatic. - core/db.py: _run_alembic_upgrade now plans the run (up_to_date / pending / unknown_revision), snapshots first when migrations will execute, and raises MigrationError on a mid-flight failure — startup stops with the backup path named instead of continuing on a half-migrated DB. The #552/#547 unknown-revision class stays non-fatal (warn + additive reconcile). - core/changelog.py + GET /api/settings/changelog: parse the shipped CHANGELOG.md (single-line and wrapped bullet styles) into structured releases. - GET /api/settings/db-backup: newest pre-migration backup for the panel. Rust (bootstrap.rs): - #314 heal guard: an exit-signature match alone can no longer delete the venv — venv_rebuild_justified requires a structural problem or a failed direct interpreter probe; a venv that probes healthy is kept and the real error surfaced. Drift/repair remains in-place `uv sync` (non-destructive). - CHANGELOG.md now ships as a bundle resource and is copied/refreshed into the project dir so the changelog endpoint works in packaged installs. Frontend (Settings → Updates): - Available update shows its actual release notes (updater metadata body) through a safe markdown-lite renderer (text nodes only, refs stay plain). - "Your data is backed up before every update" line with the latest backup timestamp from the new endpoint. - "What's new" changelog reader (accordion, newest expanded) over the shipped CHANGELOG.md; GitHub releases list reuses the same renderer. - One-time, non-blocking "What's new" footer pill after an update (persisted last-seen version; fresh installs baseline silently). - All strings via t() with en keys (other locales fall back to English). Tests: db backup/rotation/failure-path units, migration-safety units, changelog parser (both bullet styles + real CHANGELOG.md), endpoint tests, route inventory regenerated, Rust decision-logic + probe tests, vitest suites for renderer/viewer/panel/pill logic. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e2c4ea93b0 |
fix(models-settings): surface async install errors, disk-space guard, cancel wiring, honest restart (#908)
Live-audit fixes for the Models settings surface — the P1s were cases where the feature silently didn't work for the user. P1-A — Async install errors were invisible. The `install_error` SSE event carries excellent mirror-aware text (#890 core/failure.py), but the Model Store auto-purged the errored row ~800ms later (same as a success) and the first-run WizardLibrary DELETED the row without ever reading `ev.error`. The SSE→rowState reduction is now a pure, tested reducer (downloadReducer.js / reduceWizardDownloadEvent); only SUCCESS terminals auto-purge (isAutoPurgeTerminal), an error persists on the row with inline text + Retry + Dismiss (Model Store) / a Retry (wizard). P1-B — No disk-space check on install. `POST /models/install` now compares the FDL-05 plan's exact `to_download_bytes` (+ MIN_FREE_GB headroom) against `shutil.disk_usage(cache).free` BEFORE downloading and emits an actionable install_error naming the sizes (needs X, headroom Y, have Z) instead of failing mid-download. `/models` also surfaces `disk_free_gb` in the header. MIN_FREE_GB + disk_free_bytes are single-sourced in setup/models.py (wizard delegates). P2-A — Wired the orphaned cancel. `POST /models/install/cancel` (FDL-11) had zero frontend refs; the in-progress row now shows a Cancel button that calls it and transitions the row to install_cancelled. P2-B — Honest restart_required. The HF-mirror PUT returned restart_required:true unconditionally; it now returns true only when the persisted value actually changed, with accurate copy (Model Store downloads use the new mirror immediately — resolved per-call; only transformers model loads need a restart). P3 — i18n the un-localized panels (HFMirrorPanel, ApiKeysPanel source labels/help/status, MODEL_ROLE_LABEL) via new en.json keys; other locales fall back to en. Tests: new tests/test_install_disk_space.py (reject-when-over-budget incl. the worker wiring; allow-when-fits; degrade on unknown size/unprobeable volume), updated tests/test_hf_mirror_settings.py (change-only restart_required), and new frontend reducer + column-render tests for install_error persistence, Retry, Dismiss, and Cancel. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b3c18db33f |
feat(settings): Storage panel — real disk usage, category breakdown, and low-space warnings (#906)
Settings → Storage now opens with a Disk usage panel backed by a new loopback-gated GET /api/settings/storage endpoint: - Per-volume totals (grouped by st_dev) + du-style sizes for everything the app owns: the HF model cache (with its ~10 largest models), the app data dir broken into voices/outputs/dub_jobs/batch/preview/ database/logs/other subtotals, engine venvs (backend/engines/*/.venv + the app venv), and omnivoice* entries in the OS temp dir. - Bounded scanning: per-category 10 s deadline → partial totals with an "unreadable" warning instead of a hung request; results cached in-process for 5 minutes, ?refresh=1 forces a rescan; the walk runs in a worker thread so the event loop never blocks. - Server-side warnings reuse the setup wizard's MIN_FREE_GB: free < min → critical, free < 2×min → low, volume holding the cache/data >90% full → volume_pressure, unreadable/timed-out paths → unreadable. The panel renders severity-colored banners, a data-volume gauge, proportion bars per category, Open-folder buttons (existing /export/reveal pattern), a Model Store jump for reclaiming model space, and the existing clear-logs action on the logs row. A critical warning is also surfaced outside Settings via the app-wide toast — once per session. All strings via i18n (en fallback). Tests: tests/test_storage_report.py (sizes, thresholds, cache/refresh, timeout partials, endpoint wiring) + StorageUsagePanel.test.jsx (categories, banners, once-per-session toast, refresh=1, error state); route added to tests/fixtures/api_routes.txt via the dump script. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e7fc37d438 |
fix(settings): retire legacy LLM endpoint panel, surface env overrides, fix Cloudflare account + fast-fail probes (#907)
Live-audit fixes for Settings → LLM Providers / Translation.
Retire the legacy LLMEndpointPanel from the UI (backend endpoint kept).
TranslationTab no longer embeds the inline endpoint panel — it now points to
Settings → LLM Providers (openSettingsTab('llm-providers')), which fully covers
it via the `custom` provider (a lone TRANSLATE_BASE_URL still resolves to
`custom`). Kills the panel's lying "reachable" badge, its hardcoded-English
strings, and one of three duplicate TRANSLATE_* surfaces. The third duplicate —
TranslationTab's "Provider keys" collapsible — drops the TRANSLATE_* trio
(now owned by LLM Providers) and keeps only the DeepL/Microsoft translator
keys; its toast no longer claims "saved for session" (these are in
PERSISTENT_KEYS, restored at startup). GET/PUT /api/settings/llm-endpoint is
untouched (DubTab gates Cinematic off it; tests + route inventory cover it).
Surface env overrides. describe() now reports base_url_from_env / model_from_env
/ active_from_env (mirroring key_from_env). The panel disables env-pinned
base_url/model/account fields with an explainer, and — when
LLM_DEFAULT_PROVIDER pins the active provider — disables make-active and shows a
banner, instead of silently reverting the user's edit / no-oping the button.
Fix the Cloudflare account-id flow (broken two ways): describe() now returns the
stored account_id (the field no longer resets to empty) and shows the RAW
base_url template ({account_id} kept literal) instead of the substituted value;
save_overrides drops a base_url override equal to the built-in default, so the
UI posting the shown value back can't freeze the URL — later account-id changes
take effect again (also self-heals if a default URL changes in a release).
Fast-fail the Test / Fetch-models probes. Pass max_retries=0 to the probe
OpenAI clients so a 429/timeout returns in seconds instead of ~34s on the SDK's
default retry ladder. /models now returns truncated:true when capped at 200 and
the UI hint reads "first 200 shown".
Tests: registry env-flag + Cloudflare round-trip/no-freeze regressions; router
truncation + max_retries=0 assertions; panel disabled+explained + banner;
new TranslationTab test (pointer wired, legacy panel gone, TRANSLATE_* dropped).
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
83e71c5689 |
fix(asr): close the #730 residuals — chunked dub wedge shares the guarded reset; repeated timeouts recommend the crash-isolated engine (#895)
Residual A — the chunked dub-stream had a PARALLEL wedge mechanism (its own ping-loop timeout, its own _reset_pool_on_wedge, a dead-end "Try restarting the server" message). A wedged chunk now routes through the SAME run_transcribe_guarded bound+reset as the whole-file paths (#731/#851): the guard resets the poisoned pool once per wedged attempt (no double-reset on retry) and the user sees the actionable ASRTimeoutError. The reset logic is extracted to asr_backend.reset_pool_after_wedge — one shared mechanism, so the semantics can't drift again. run_transcribe_guarded also gains a timeout_env param so chunk errors name OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S instead of the whole-file knob. Residual B — the crash-isolated ASR sidecar (#393, faster-whisper-isolated) is wired as an explicit ESCAPE HATCH, not a default: - selectable end-to-end: Settings engine list gets an explanatory install_hint; honest gpu_compat ("cuda","cpu" — it wraps the same CTranslate2 engine as faster-whisper); get_active_asr_backend now hands back a process-wide singleton for subprocess-isolated backends (a fresh instance per request would leak atexit hooks and respawn the sidecar — reloading its model — on every transcribe). - on the SECOND consecutive guarded timeout-with-reset in one session (resets aren't recovering the hang; the wedged thread keeps its VRAM), the error the user sees + the log recommend switching to the isolated engine in Settings → Engines. Never auto-switched (owner rule: no silent behavior divergence); a completed transcribe resets the streak. Tests (fail-before/pass-after verified against origin/main): wedged-chunk SSE integration (reset count + actionable error + recommendation surfaces), consecutive-timeout streak (fires at 2, resets on success, suppressed when already on the isolated engine), timeout_env parametrization, shared-reset helper, isolated backend in list_backends with hint + honest availability, singleton caching, gpu_compat matrix entry. Docs: troubleshooting §14 gains the chunk knob + escape-hatch guidance. Closes the residuals tracked on #730. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6e600c48cb |
fix(generation): classify network/download failures — stop mislabeling every unknown error as OOM (#880) (#893)
A kittentts first-use HuggingFace download died with httpx's "Cannot send a request, as the client has been closed", and the generation error classifier's catch-all fallback told the user (CPU-only ~80 MB ONNX engine, 12 GB-VRAM box) they were OUT OF MEMORY and to press Flush — the wrong remedy for a network failure. Three-part class fix: - generation.py: new #880 branch (before the OOM hint) classifies httpx/requests transport failures — matched over the whole exception chain (type names like ConnectError/ReadTimeout plus stringified signatures like "client has been closed") — as a download/network problem with a retry/check-connection remedy. - generation.py (the real class bug): the OOM hint is no longer the catch-all. It now requires an actual OOM signature (typed OutOfMemoryError/MemoryError anywhere in the chain, or CUDA/MPS/CPU allocator wording); genuinely unknown errors surface as unrecognized with the underlying detail instead of a false "ran out of memory". - tts_backend.py: KittenTTS's first-use load retries exactly once with a fresh HF Hub client (huggingface_hub.utils.close_session()) on the specific closed-client failure — hub ≥1.x shares one global httpx client, and a closed one is recoverable, so the download self-heals instead of failing the generation. Fail-before/pass-after tests: classifier (closed-client message, wrapped httpx type names, unknown error, real OOM signatures incl. typed OutOfMemoryError, WinError 1455) + the retry helper (recovers once, walks the chain, no retry on unrelated errors, single-shot). Fixes #880 Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
14f1257d1f |
fix(errors): name the configured HF mirror when a model download fails (#874) (#890)
When a non-default HF_ENDPOINT (Settings → Models → Hugging Face mirror,
e.g. hf-mirror.com) is configured and a model load/download fails with a
connectivity error, the raw transformers message ("We couldn't connect to
'https://hf-mirror.com' to load the files…") leaked to the UI as a bare 500
with no next step.
Class fix — one shared classifier in core/failure.py covers every surface:
- classify()/build_failure(): new HF_MIRROR_UNREACHABLE class with a dynamic
hint that names the configured mirror, says it may be down, points at
Settings → Models → Hugging Face mirror, suggests the official endpoint
when the model isn't cached, and notes the restart requirement (HF reads
HF_ENDPOINT at backend start). Checked before the video-download network
class so a model download's "timed out" no longer gets the "video server"
hint. Feeds /model/status and every build_failure event (dub, tasks).
- main.py global 500 handler: appends the hint to the surfaced detail, so
ALL routes that can leak a model-load error benefit (generate, dub,
archetypes, …), not just TTS generate.
- setup/download.py install SSE: the install_error event gets the same hint.
- error_journal: "couldn't connect to" / "max retries exceeded" now classify
as NETWORK_ERROR (was UNKNOWN) for auto-attached bug reports.
- model_manager (#886 family): the "cache incomplete and could not be
auto-repaired" message now names WHY the auto-repair failed (mirror
outage, offline mode, full disk no longer read identically), which also
lets the mirror hint fire on that surface when applicable.
Fail-before/pass-after regression tests in tests/test_hf_mirror_error_class.py
(12 of 13 fail on main).
Fixes #874
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d58010fe1b |
feat(dictation): rebuild to Wispr-Flow quality — live waveform, streaming commits, honest insertion, polished text (#888)
* feat(dictation): rebuild to instant-feedback quality — waveform, streaming commits, honest insertion, text polish Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): dictation rebuild entry Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(lint): Array.from over new Array(n) — oxlint no-array-constructor Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
da9315815d |
feat(settings): LLM provider testing pass — latency + classified errors, model discovery, full i18n, router tests (#887)
Co-authored-by: mergetest <test@local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f4e318f9f2 |
fix(translate+dub): wire Cinematic/Autofit to the LLM Providers registry; retry a wedged transcribe chunk instead of dropping it (#867)
Two bugs from real reports: 1. LLM not wired — translator._llm_client()/_llm_model() read TRANSLATE_*/OPENAI_* directly, bypassing the LLM Providers registry (#854). So a provider set up in Settings → LLM Providers never powered Cinematic/Autofit. Now resolves the ACTIVE provider (base_url/key/model) via llm_providers; the 'custom' provider still maps TRANSLATE_* so legacy env setups keep working. 2. Transcription 'missing the beginning' — the chunked dub transcribe dropped a whole chunk's window on failure/timeout (returned empty segments, no retry). A transient wedge on the FIRST chunk (whisperx cold-loads its model there, the #730 hang) therefore lost the start and left only middle+end. Now retries a failed/timed-out chunk once on a fresh pool (OMNIVOICE_TRANSCRIBE_CHUNK_ATTEMPTS, default 2) so the recovered chunk fills the hole. Imports + dub_transcribe/translator/llm_providers tests green. Co-authored-by: mergetest <test@local> |
||
|
|
29269b9cf0 |
feat: LLM Providers page + Autofit translation quality (fit-to-segment-time) (#838) (#854)
* feat(llm): multi-provider LLM registry + encrypted key storage + settings API (v0.3.8, phase 1)
Foundation for the LLM Providers settings page and timing-aware (Autofit)
translation. Every provider in the shipped .env is OpenAI-compatible, so one
client drives all of them via a registry instead of a class-per-provider.
- llm_providers.py: registry of 16 providers (OpenAI, OpenRouter, Groq,
Cerebras, Google AI, Mistral, Cohere, NVIDIA, GitHub Models, Cloudflare,
HuggingFace, SambaNova, SiliconFlow, + local Ollama/LM Studio + Custom).
Field resolution precedence env → encrypted store → default; active-provider
selection (LLM_DEFAULT_PROVIDER → stored → first keyed remote; local requires
explicit pick so we never assume a local server is up). Legacy TRANSLATE_*
maps to the Custom provider (keyless-with-base_url preserved).
- settings_store.py: generic ENCRYPTED secrets (get/set/clear_secret,
list_secret_names) reusing the HF-token Fernet path; get_text/set_text now
refuse the secret namespace (no ciphertext leak).
- llm_backend.py: OpenAICompatBackend resolves the active provider's
base_url/key/model from the registry. Backward-compatible.
- settings API: GET /llm-providers, PUT /llm-providers/{id} (encrypted key +
overrides), POST /llm-providers/active, POST /llm-providers/{id}/test.
Loopback-gated; never returns key material.
- 11 registry tests; existing llm-endpoint/openai-available tests still green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(settings): LLM Providers page — configure any provider's key/URL/model + Test + set active (v0.3.8, phase 2)
New Settings → System → LLM Providers pane (Brain icon, searchable). Lists all
16 registry providers; pick one to configure its encrypted API key, base URL,
model (and Cloudflare account id), Test the connection with one round-trip, and
'Save & use for translation' to make it the active provider for Cinematic/
Autofit. Keys are write-only from the UI (masked placeholder, never echoed);
env-set keys show as read-only. Local providers (Ollama/LM Studio) need no key.
- LLMProvidersPanel.jsx: provider selector + per-provider config + Test/activate,
following the LLMEndpointPanel pattern (apiJson/apiFetch/apiPost, SettingsSection
primitives).
- settingsCategories.jsx: new 'llm-providers' category under System + Brain icon.
- Settings.jsx: route the category to the panel.
- en.json: settings.llm_providers label.
- Frontend build passes.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(translate): Autofit quality style + one-click LLM setup from the dub menu (v0.3.8, phases 3-4)
Autofit = Cinematic + a strict 'never exceed the segment time' fit. The LLM
rewrites each translated line so its target-language reading time fits within
the slot, preserving the video timing without harsh audio time-stretch.
Backend:
- speech_rate.adjust_for_slot(strict=): strict caps the accepted upper ratio at
1.0 (fit within slot) vs Cinematic's 1.08; best-effort, degrades gracefully
with no LLM.
- dub_translate: quality='autofit' takes the LLM refine path and runs the fit
pass with strict=True; reports quality_used accurately.
- TranslateRequest.quality doc note.
Frontend:
- 'autofit' added to the quality control (Settings Translation + dub menu) and
the TranslateQuality type.
- Dub menu: picking Cinematic/Autofit with no LLM no longer dead-ends on a toast
— it offers a one-click 'Set up' that routes to Settings → LLM Providers,
with copy about fitting translations to segment time (#838).
- 4 strict-fit tests; frontend build green; i18n keys added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(translate): document Autofit quality + the LLM Providers page (v0.3.8, phase 5)
- CHANGELOG [0.3.8] Added: Autofit style + LLM Providers page.
- docs/dubbing/translation-engines.md: Fast/Autofit/Cinematic quality section
and an LLM Providers setup section (16 providers, encrypted keys, offline
Ollama/LM Studio, env overrides).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(api): add /api/settings/llm-providers routes to the route-inventory snapshot
Regenerated tests/fixtures/api_routes.txt for the 4 new LLM-providers endpoints
so test_route_inventory_matches_snapshot passes (keep-main-green).
* style(frontend): oxfmt the LLM Providers panel + dub quality control (format:check green)
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
c5c57508b3 |
fix(device): fall back to CPU when the GPU arch is unsupported, not 500 every generate (#756) (#757)
* fix(settings): contain + tighten the whole Settings surface (measure cap, container-query stacking, wrap the shared rows) Two systemic issues drove 'too spread out' + 'elements go out of view' across many Settings pages: 1. Spread — .settings-content capped at 1280px, so on wide windows every label-left/control-right row left a huge void. Introduce a --settings-measure token (720px, macOS-like) + --settings-rail, and cap the content to it, left-aligned under the nav. One token now controls the reading width. 2. Overflow + bad responsiveness — the row stack break was a *viewport* media query (560px), but the 168px nav rail means a 760px-viewport window only has ~530px of content, so rows went side-by-side in a cramped box. Make .settings-content a container (container-type: inline-size) and stack on the CONTENT width via @container, keeping the viewport @media as a fallback for the .st-row instances used outside Settings (Splash/FirstRun/Dub/SetupWizard). 3. The shared .perfpanel__row (button/badge row reused by 6+ panels: RemoteBackend, HFMirror, LLMEndpoint, Pronunciation, MCPBindings, …) was an inline-flex with no wrap and no max-width, so it ran off the right edge — add flex-wrap + max-width:100% + min-width:0. Plus two rigid-width fixes that escaped the row cap: ApiKeys input min-width:220→0, Appearance scale floor. Frontend builds clean; tokens, @container query, and the wrap all verified in the emitted CSS bundle. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(settings): center the settings block + tighten measure (kill the lopsided right void) The capped content was left-aligned, so on a wide window everything jammed to the left with a dead empty third on the right (screenshot). Center the whole settings block (nav rail + content) as a unit via max-width + margin-inline:auto, and drop the measure 720→660 so label→control rows read denser. The cap is computed from the tokens (rail + gap + measure + page padding) so the content track lands exactly at --settings-measure. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(device): fall back to CPU when the GPU arch is unsupported, instead of 500-ing every generate (#756) get_best_device() called check_device_compatibility() and, on an unsupported compute capability, only LOGGED a warning then still returned 'cuda' — so the model loaded on a GPU whose kernels can't launch and every generate 500'd with 'CUDA error: no kernel image is available for execution'. Both a too-old card (Pascal sm_61, GTX 10-series) and a too-new one (Blackwell sm_120 on pre-cu128 wheels) hit this. Now an unsupported arch falls back to CPU (works, just slower) with a clear warning; OMNIVOICE_FORCE_CUDA=1 overrides. Belt-and-suspenders: _oom_friendly_reraise classifies a raw 'no kernel image is available' as an unsupported-GPU error (switch to CPU / install matching torch) rather than the OOM/Flush message. Tests: get_best_device → cpu on incompatible, stays cuda on compatible, honors the force override; reraise gives the actionable GPU message, not OOM. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(device): patch detect_host_caps via string path so the #756 fallback test is full-suite robust The first version aliased the import + inserted backend on sys.path, which patched a module copy get_best_device's local 'from core.device_caps import detect_host_caps' didn't resolve in the full suite (passed alone, failed in CI). Use the string-form monkeypatch target; verified passing alongside the other device/model tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): fold #757 device-fallback entry into [0.3.8]; drop the merge's stale [Unreleased] dupe --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e347f99542 |
fix(tts): bound + reset the GPU pool on a hung generate so it can't brick the backend (#730 class) (#851)
* fix(tts): bound + reset the GPU pool on a hung generate so it can't brick the backend (#730 class) A GPU job that wedges on some Windows+CUDA setups occupies its worker forever — run_in_executor can't cancel the thread — so on the 1–2 worker pools we ship, one stuck job starves every other request and the next action surfaces as the misleading "Can't reach the local backend" even though the process is alive. ASR/dub/model-load already bound+reset the pool on hang (#730). The TTS **generate** paths (generation.py, tts_stream.py) were the last unguarded GPU dispatch — and the residual on-main reports (#850 #802 #755 #723 #721, plus the 0.3.7 generate cohort) all fail on generate:start (audio). - model_manager: add run_on_gpu_pool_guarded() + GpuJobTimeoutError, a generalized version of the ASR guard so every GPU dispatch shares one bound+reset recovery path. Env-tunable via OMNIVOICE_GENERATE_TIMEOUT_S (default 300s). - generation.py: route both inference branches + the reference-clip transcribe through the guard; map a timeout to an actionable 503. - tts_stream.py: same guard on the streaming path (timeout → error frame). - test_generate_timeout_730: fail-before/pass-after regression (timeout resets pool + restores capacity, happy path, env override, no-reset exec). - docs + CHANGELOG: extend troubleshooting §14 to cover generate; document the new env var. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(tts): extend the GPU-pool hang guard to batch/dub/archetype/openai-compat generate (#730 class) The generate-hang class wasn't only in Studio + streaming: batch generate, the dub per-segment + preview generate, archetype preview render, and the OpenAI-compat /v1/audio/speech path all dispatched the TTS model to the GPU pool with no wall-clock bound either. Any one of them wedging on a Windows+CUDA hang starves the pool and bricks the backend the same way. Route all of them through run_on_gpu_pool_guarded so the whole class is closed — a hung generate anywhere resets the pool and returns an actionable timeout instead of a dead backend. Batch/dub recover per-segment on a fresh worker; drop the now-dead loop/_gpu_pool/asyncio locals ruff flagged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
522bbddccf |
feat(translate): highlighted Install affordance for uninstalled engines + dismissable/auto-clearing error banner (#847)
Two related Dub-tab translation-flow fixes, one PR. TASK 1 — proactive, highlighted Install affordance in the translate engine selector (replaces "find out only via a translate-time 400"): - FROM-SOURCE lane (activeEngineUnavailable && !enginesSandboxed): the muted install chip is promoted to a HIGHLIGHTED brand-accent Install button, still wired to handleInstallEngine(translateProvider) with the installing/disabled state. Selecting any uninstalled engine surfaces it immediately. - FROZEN lane (enginesSandboxed): pip install is impossible in the read-only, signed packaged env, so the disabled "needs dev install" span becomes an equally highlighted button opening a popover with (1) the exact install command + copy-to-clipboard, (2) one-click "Switch to Argos (bundled, offline)" — the guaranteed importable escape hatch, and (3) a Docs link via the existing Tauri shell.open path. Gated on the existing `sandboxed` flag, not platform. - Single-source install command: new translation_engines.install_command() is the one source of truth; list_engines() stamps `install_command` per engine and BOTH the argos + deep_translator translate-time 400 messages build their command from it, so the proactive button and the 400 can't drift. engines.ts gains `install_command: string | null`. TASK 2 — the translation error banner now dismisses and clears (class fix): - Root cause: handleTranslateAll never cleared dubError, so a stale 400 survived even a successful retry. It now clears at the start of every attempt. - Corrective-action clears (whole class): changing the engine and installing the package both clear dubError (wrapped setTranslateProvider + handleInstallEngine in DubTab). - DubFooter's banner gains a × dismiss and a guarded auto-timeout (skipped while generating so live per-segment errors persist). i18n: 8 new dub.* keys translated across all 21 locales. Docs: new docs/dubbing/translation-engines.md (from-source vs packaged build) linked from the popover Docs button + a troubleshooting cross-reference. Tests: FE regression for both lanes + never-installs-when-sandboxed + banner dismiss/auto-clear; BE regression that list_engines() install_command is embedded verbatim in the dub_translate 400s. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cc95f526e0 |
fix(asr): reset the GPU pool when a chunked dub-stream chunk wedges too (#730) (#742)
The whole-file transcribe paths recover from a wedged worker via run_transcribe_guarded's pool reset (#731), but the chunked dub transcribe-stream only recorded a per-chunk timeout error and moved on — leaving the stuck thread holding its GPU-pool worker, so subsequent chunks / a concurrent TTS generate could still starve into 'can't reach backend'. Reset the pool on the per-chunk TimeoutError via a small _reset_pool_on_wedge() helper (best-effort, no-op for a plain executor). Closes the residual on #730. Tests: helper resets a reset-capable pool and no-ops a plain ThreadPoolExecutor. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
0fc9f2afec |
fix(asr): bound every transcribe path + reset the GPU pool on hang so a wedged ASR can't brick the backend (#730) (#731)
A whisperx/CTranslate2 transcribe can hang hard on some Windows+CUDA setups and never return. ASR shares the small (1-2 worker) _gpu_pool with TTS, so one stuck worker starved every other request — the next TTS generate then surfaced as "Can't reach the local backend" though the process was alive (#720/#721/#723). Two parts: - Bound the three remaining unguarded whole-file transcribe paths (dub whole-file dub_core.py, batch.py, live-dictation capture_ws.py) with run_transcribe_guarded, matching the dub-QC/dictation/OpenAI paths that were already bounded by #656. - On timeout, run_transcribe_guarded now calls executor.reset() when the pool supports it (_ResilientGpuPool, already built for the model-load-timeout case in #589/#599): the wedged worker is abandoned and the next submit gets a fresh one, restoring capacity without an app restart. Best-effort — a plain ThreadPoolExecutor (tests) just gets the bound + actionable error. Regression tests: pool.reset() is invoked on timeout; a non-reset pool still bounds cleanly. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a9948de99e |
fix(generate): classify [Errno 32] Broken pipe as a lost-pipe error, not OOM (#715) (#722)
A BrokenPipeError surfacing from generation means the backend's stdout/stderr pipe to the desktop shell that launched it closed mid-render (an orphaned or relaunched backend) — not out of memory. _oom_friendly_reraise mislabeled it "ran out of memory — try Flush," which never helps. Add a BrokenPipeError / [Errno 32] branch (same pattern as the #705 WinError-193 and #437 permission branches) that tells the user to restart the app instead. main.py already wraps sys.stdout/stderr to swallow EPIPE; this catches the C-level writes inside the native engine/torch that escape that guard. Regression test covers both the typed BrokenPipeError and a string-wrapped "[Errno 32] Broken pipe". Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ea4d5e7839 |
fix(generate): self-heal schema + don't 500 a generated clip on a history-write fail (#710) (#714)
A synth that already produced and saved its audio could still return a 500: 'no such table: generation_history' — a DB that somehow missed schema init (init_db's executescript never took) made the history INSERT raise after the clip was done, losing the user's generation to a logging side-effect. - Add db.ensure_schema(): idempotent CREATE ... IF NOT EXISTS + additive column reconcile (no _migrate/alembic), safe to call from a write path. - Generation history write now self-heals: on a sqlite OperationalError it runs ensure_schema() and retries once; if it still fails it logs and returns the audio anyway. A history-logging failure can never fail the generation. Regression test: the write raises 'no such table: generation_history' before the heal and succeeds after (fail-before/pass-after), plus ensure_schema idempotency. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2c2e493df8 |
fix(dub): stream segments to disk to stop long-video RAM spikes (#639) (#709)
Takes over and completes #639 (original work by @trungthanh1288). Dub generation held every segment's audio in RAM until final mix, so long/feature-length dubs and big batches could exhaust memory. Segments now stream to disk as rendered; the final track assembles from those files via a 30s-chunk memmap writer, so peak memory stays flat regardless of length. Completed on top of the original PR: - Watermarking: keep the project's 'every OmniVoice audio carries the signature' guarantee without double-marking. Since seg_<id>.wav is BOTH the downloadable file AND the assembly input, mark each fresh segment once at synthesis and drop the per-chunk embed in the memmap writer (the final mix inherits the mark) — main's proven policy. Verified with real AudioSeal: 0.9999 detect confidence on the final track and on seg WAVs; cached/silence not re-marked. - Fix a crash regression: zero/negative-duration segments returned an in-memory zero-length entry instead of writing empty audio (which raised). Regression test added. - Perf: drop per-segment gc.collect(); throttle empty_cache() to every 16th call (the replaced code batched I/O to keep this off the hot path). - Clean up the mix_<id> temp WAVs after assembly. - Rewrite the watermark test for the multi-chunk (>30s) path; assert both the final track and the seg WAV are marked, with no double-mark. 212 passed / 1 skipped; route inventory clean. Co-authored-by: mergetest <test@local> Co-authored-by: trungthanh1288 <trungthanh1288@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
aa17e3319f |
fix(model): route every OMNIVOICE_MODEL read through the resolver; tighten WinError 193 match (#693, #705) (#707)
Follow-up from independent verification of #693/#705. #693 (whole-class): the resolver only guarded the model-load site. A leaked engine id in OMNIVOICE_MODEL still hit four other raw reads — most importantly preload_model()'s model_info() probe, which failed on the bad value and SILENTLY disabled warm-up (first /generate then ate the full load). Plus the Settings 'model_checkpoint' display, the loaded-models list, and the engine_id baked into exported persona bundles. Route all of them through resolve_omnivoice_checkpoint() (personas keeps its '' unset marker, sanitizing only a set value). Add a source-level recurrence guard so a future raw read can't reintroduce the class. #705: tighten 'winerror 193' -> '[winerror 193]' so the substring can't also match WinError 1930-1939 (the portable 'is not a valid win32 application' clause still covers non-Windows formatting). 48 tests pass (resolver + guard + audio-guard + route inventory); edited routers/services import clean. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
883a06e9c0 |
fix(generate): classify WinError 193 as a corrupt native component, not OOM (#705) (#706)
A synth failure from a corrupt or wrong-architecture native binary on Windows
([WinError 193] %1 is not a valid Win32 application — torch, ffmpeg, or a
bundled engine binary) fell through to the generic OOM message ('ran out of
memory — try Flush'), sending the user down a path that can't help.
_oom_friendly_reraise() now detects the WinError 193 / 'is not a valid Win32
application' signature (before the OOM fallback, joining the existing
torch.compile / decode-glitch / bad-instruct cases) and surfaces an actionable
'reinstall or repair that component; Flush won't help' message. Regression test
added.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ee638bda6c |
feat(tts): user pronunciation dictionary (expressive-tts slice 1) (#685)
* feat(tts): user pronunciation dictionary (expressive-tts slice 1) Per-term, per-language pronunciation overrides applied to text before synthesis, so names, brands, and acronyms come out right across generate, longform, and dub. Closes part of the #1 perceived-quality gap vs ElevenLabs (pronunciation dictionaries). First slice of docs/specs/01-expressive-tts.md. - Schema: additive `pronunciation_entries` table (alembic 0008, mirrored into _BASE_SCHEMA; tested upgrade — idempotent, downgrade, converge, back-compat). - Service: extend pronunciation.py to load enabled entries (cached) and apply longest-first, word-boundary-aware, per-language (global '*' + lang match, lang overrides global), reusing the existing ReDoS-safe matcher. - Inline one-off `[[term|replacement]]` overrides that don't persist and don't collide with [voice:]/[pause]/[Name]/SSML-lite (resolved pre-chunking). - API: /pronunciation CRUD + /test dry-run + import/export (loopback-guarded). - Apply point: generation.py after language resolves, before chunking — covers native + pluggable engines. - UI: PronunciationPanel in Settings → General; all strings via i18n. - Tests: migration lifecycle, CRUD, per-language, precedence, inline override, apply-at-synth. Route snapshot regenerated (+7). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(security): bound inline-override regex (ReDoS) + annotate parameterized UPDATE CodeQL flagged py/polynomial-redos on the [[...]] inline-override regex: [^\]] also matches [, so an unterminated run of [ allowed O(n) rescans from O(n) positions. Bound the inner class to {0,256} (linear; an inline override is a short respelling). Annotate the dynamic UPDATE (B608) — its column fragments are fixed literals and every value is a bound parameter; not an injection vector. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
022a3bd6b9 |
feat(dictation): live local dictation via sherpa-onnx + Voice settings panel (#683)
* feat(dictation): live local dictation via sherpa-onnx + Voice settings panel Add a sherpa-onnx ASR engine alongside the existing Whisper/NeMo dictation path, powering a genuinely live experience: as you speak, words type straight into the focused field (streaming partials via a new simulate_type command, self-correcting with backspaces) and commit per pause. Backend: - SherpaDictationBackend + sherpa_dictation registry of the 7 models (Parakeet TDT v3/v2, streaming Zipformer EN/ZH/bilingual, Paraformer bilingual, Whisper Tiny) from csukuangfj/* int8 HF repos; CPU provider for cross-platform parity. - /dictation/models + /dictation/prefs router; get_capture_asr_backend() honors the selected dictation model. get_active_asr_backend() (dub transcription) and the legacy WebM/Opus capture path are untouched. - True streaming over /ws/transcribe (OnlineRecognizer: live partials + per-endpoint finals); offline models surface partials via short re-decode. Frontend: - New "Voice" settings panel (enable, Toggle/Hold mode, model picker with offline/streaming/recommended badges + per-model download/delete). - Live word-by-word typing via simulate_type (enigo) with prefix-diff delta and backspace correction; paste fallback retained, no double-insertion. Deps: sherpa-onnx>=1.13.3 (+ sherpa-onnx-core); uv.lock regenerated, Docker frozen-install verified. API route-inventory snapshot updated. 40+ new tests. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dictation): register sherpa-onnx-asr engine in README + features inventory Fixes the docs-drift CI guard: the new sherpa-onnx-asr ASR engine existed in the registry but not in docs/features.yaml or README. Adds the live-dictation engine row to the ASR Engines table, bumps the engine counts (8→9), and adds the inventory entry. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(changelog): fold live-dictation into the [0.3.8] section main is 0.3.8 (untagged), so the dictation feature belongs in that release, not a separate [Unreleased] block. Merge the two Added lists under one [0.3.8], refresh the headline to lead with live dictation, and correct the capture description to reflect live word-by-word typing (not paste-on-pause). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b7cecde57e |
feat(setup): faster downloads by default + prominent, encouraged HF-token entry (#669)
Two changes that make first-run downloads faster and easier to speed up further.
1. Segmented (multi-connection) downloader is now ON by default. The app forces
the legacy-LFS path (HF_HUB_DISABLE_XET=1) for clear progress, but that path
is single-stream and slow — which is why downloads felt sluggish. The built-in
IDM/uGet-style segmented accelerator (parallel byte-ranges, live speed/ETA)
was already implemented but defaulted OFF. Flip it ON: it only engages when
Xet is inactive (the default), and ANY failure falls back to snapshot_download
("can never compromise a correct install"). Pure-httpx, cross-platform,
auth-safe (token never forwarded to a CDN). Override with
OMNIVOICE_SEGMENTED_DOWNLOAD=0.
2. The Hugging Face token field is now a prominent, always-visible card right
above Continue — was a collapsed "advanced" fold almost nobody opened. A free
token gives authenticated downloads (higher rate limits, fewer stalls), so it
pairs with change #1 to keep the parallel fetch from getting throttled. The
card leads with the speed benefit, shows a saved-state, and adds a one-click
"Get one free →" link to huggingface.co/settings/tokens.
Docs: downloading-models.md updated — the legacy-LFS section now documents the
default-on segmented accelerator + the HF-token speed tip, and the tuning table
reflects OMNIVOICE_SEGMENTED_DOWNLOAD=0 as the disable knob (docs-sync).
Test: test_segmented_download_default.py pins the new default ON and that the
env override still disables it; existing FDL-08 behavior tests stay green.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
81a007552b |
fix(generate): classify a bad-instruct error as a 400, not a 500 "ran out of memory" (#664) (#665)
A user typed free-form prose ("Speak with high energy … like a podcast host")
into the voice-design instruct field and got a **500** whose message read "TTS
engine stopped mid-generation. This usually means it ran out of memory. Try the
Flush button …" — with the real cause ("Unsupported instruct items found …")
buried as the underlying error. The user is told to Flush for an OOM that never
happened; the actual problem is a rejected instruct.
Root cause: `_resolve_instruct` raises on unknown/conflicting instruct items, but
by the time the error reaches `_oom_friendly_reraise` it's no longer a bare
`ValueError` (a lower layer wraps it), so the route's `except ValueError -> 400`
guard misses it and it falls through to the generic OOM `RuntimeError`. v0.3.7
has had that guard since v0.3.6 yet still produced the OOM message — proving the
error arrives wrapped, so type-based detection is insufficient.
Fix: in `_oom_friendly_reraise`, detect the instruct-validation **message
signature** ("unsupported instruct items" / "conflicting instruct items" / "in a
single instruct") regardless of exception type and re-raise a clean `ValueError`,
so the route returns a **400 with the instruct guidance** instead of a 500 OOM.
This is version-independent and complements the client-side guard (#658/#612):
it also covers API/MCP callers and stored profiles whose instruct slips through.
Test: two cases in test_generation_audio_guard.py — a bare instruct `ValueError`
and one wrapped in a `RuntimeError` both reclassify to a ValueError without the
"ran out of memory" text; the generic OOM path is unchanged.
Closes #664
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
252f0d4fac |
fix(asr): bound whole-file transcription so a stall isn't reported as "can't reach backend" (#656)
A Windows/CUDA user (Vietnam) hit "Can't reach the local backend" only when dubbing/transcribing. Their log proves the backend started fine — model loaded, preload complete, 25 models — and the log ends right after `whisperx transcribing …tmp.wav`. The backend was alive; the *transcription* stalled (large-v3 ASR contending with the resident TTS model for VRAM on an 8 GB-class GPU), which the UI surfaces as an unreachable backend. Root cause (class, not instance): the chunked dub pipeline already bounds each chunk (OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S), but the *whole-file* transcribe paths ran unbounded: - dub QC re-transcribe (dub_export) - dictation (capture) - OpenAI-compat /audio/transcriptions A slow/stuck transcribe on any of these hung the request AND held a GPU-pool worker — indistinguishable from a dead backend. Fix: add run_transcribe_guarded() in services/asr_backend.py — a shared asyncio.wait_for wrapper (ASRTimeoutError, a TimeoutError subclass) with a generous env-tunable bound (OMNIVOICE_ASR_TRANSCRIBE_TIMEOUT_S, default 300 s). On timeout the request returns 504 with actionable guidance (backend is alive; free VRAM / pick a smaller ASR model / use CPU; restart to clear the stuck worker) instead of hanging forever. Wired into all three whole-file paths. Docs: new troubleshooting §14 — "Can't reach the local backend during transcription/dubbing" — explains it's ASR weight/VRAM pressure, not a network/ mirror problem, and corrects the misconception that a "Network → Restricted/Global mirror" Settings toggle exists (the Network control is LAN sharing). Serves the #602/#585/#567 "can't reach backend" cluster. Test: backend/tests/test_asr_transcribe_timeout.py — slow fn raises ASRTimeoutError with the actionable message, fast fn passes through, subclass-of-TimeoutError so the openai_compat broad catch still maps to 504. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b15acbaae9 |
fix(dub+generate): yt-dlp 403 player-client fallback (#625) + non-finite audio guard (#629) (#635)
Two independent fixes from issue triage; no version bump. #625 — yt-dlp 403 on the media download (some videos serve formats signature-protected to the default player client) is not transient, so the existing broken-pipe retry (#579) kept 403ing. The URL download now escalates the YouTube player client (tv → android → web_safari) on a 403 before giving up; a 403 no longer counts against the transient-retry budget. #629 — a numerical glitch in the model (seen on MPS) could leave NaN/inf samples that write an unreadable WAV; a downstream decode then failed with an opaque "ffmpeg returned error code: 183 / Invalid data", surfaced to the user as a misleading "ran out of memory". Sanitize non-finite samples to silence in _apply_effect_chain (single chokepoint, covers the raw path too) so the WAV is always decodable, and classify a decode/ffmpeg failure as unreadable-audio rather than OOM in _oom_friendly_reraise. Tests: 403 escalation order + success-on-alternate-client; NaN/inf sanitize + finite-passthrough + decode-error classification. Full suite 1851 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a63c8e851b |
fix(dub): speaker-aware re-split so merged speaker turns separate (#486) (#616)
Segmentation groups words into sentences BEFORE diarization, so a two-speaker exchange can land in one segment; assign_speakers_* then only relabels it with the majority speaker, losing the turn boundary (the second half of #486 — the per-speaker voice auto-assign was fixed in #490). Add a post-diarization pass that re-splits any segment whose words span >1 speaker at the word-level boundary, assigning each piece its speaker: - backend/services/segmentation.py: resplit_segments_by_diarization / resplit_segments_by_turns + a pure _resplit_core. Single-speaker segments are returned BYTE-FOR-BYTE UNCHANGED (same dict/id/text/start/end) — the no-single-speaker-regression guarantee. Pieces keep the segment's outer start/end (preserving onset-snap) and use word times for interior splits, so they exactly cover the original span. A lone mis-attributed word is smoothed, not split (diarization noise). - backend/api/routers/dub_core.py: accumulate global-timeline words alongside segments; apply the re-split after both the pyannote and FunASR-turns assign. Heuristic fallback (no word-speaker data) is untouched. 8 regression tests pin the invariant + the split/3-way/noise-smoothing/label behaviour. Full suite: 1836 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1fe68ba11e |
fix(setup): weight-aware install-state so truncated model cache isn't read as installed (#622) (#626)
A first-run user whose model download was interrupted after the config/ tokenizer files landed but before the weight shard got stranded on the Models & Engines page: GET /models computed "installed" purely from cache size on disk, so a size-positive-but-weight-less cache reported installed=true, the wizard hid the re-download button, and the model manager (Settings → Models) that could repair it was unreachable behind the wizard gate. Make install-state weight-aware. The boolean weight-floor scan now lives in models.py (the lowest module in the setup import graph) as snapshot_has_weights() + cache_is_complete(); list_models() and recommendations() downgrade a truncated cache to installed=false (+ an explicit incomplete=true on /models), so the existing "install" action re-appears and the user can re-download in-wizard. Fixes the whole class, not just /models: download.py's install-time validator now delegates to the same shared scan (one source of the floors, can't drift), matching the load-time repair in model_manager.py (#581/#606). config_only repos (pyannote/speaker-diarization-3.1 — a pipeline whose real weights live in referenced sub-repos and whose own cache is legitimately tiny) carry a new config_only:true hint in models.yaml and are exempt, so they're not false-flagged as incomplete. Tests: tests/test_mm2_lifecycle.py — snapshot_has_weights truncated-vs-complete, cache_is_complete on a truncated weight repo + config-only exemption, and list_models downgrading a size-positive truncated cache to installed=false / incomplete=true. Full backend suite green (1832 passed). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1575baca36 |
fix(design): heal validator-rejecting instruct on design voices (#594/#571/#596) (#600)
* fix(design): heal validator-rejecting instruct on design voices (#594/#571/#596) A designed voice could persist an `instruct` the engine validator rejects — either the literal "[object Object]" from a pre-fix build (#550) or freeform prose typed into the style field — so every Generate/Dub that used the voice failed with `Unsupported instruct items found in …` (400/500, and "Can't reach the local backend" when it tore down mid-render). Migration 0006 only *blanked* "[object Object]", which silently discarded the design — an Indonesian female voice then rendered male (#594). Fix the whole class by healing at every seam and rebuilding from the authoritative source (the design's saved `vd_states` category picks): - omnivoice/utils/voice_design.py: add sanitize_instruct / instruct_from_vd_states / heal_design_instruct — forgiving (never raise), drop poison/prose to valid tags, and rebuild tags from vd_states when the stored value is unusable. - profiles.py: sanitize + rebuild at save (POST) and sanitize at edit (PUT), so no poisoned instruct can ever be persisted again. - generation.py + dub_generate.py: heal whenever a profile drives synthesis, so legacy poisoned rows resolve to valid tags instead of 400-ing. - migration 0007: heal existing profiles in place (recovers gender/age/pitch from vd_states), self-contained (frozen vocab snapshot) so it never drags torch into startup; supersedes 0006's blanking. Backward-compatible. Tests: unit coverage for the healer, a migration test driving 0006->0007 on the real schema, a parity guard so the frozen snapshot can't drift, and two API guards. Corrected one existing test that had encoded the #594 behaviour. Resolves #571, #594, #596; removes a major driver of the "Can't reach backend" reports. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(cjk): allowlist migration 0007's frozen dialect-tag snapshot (#564) The 0007 instruct-heal migration carries a frozen copy of the design-tag whitelist (incl. Chinese dialect tags) so it stays self-contained; add it to the hardcoded-CJK allowlist like omnivoice/utils/voice_design.py. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
31ba6d3d27 |
fix(transcribe): surface the real ASR-load failure instead of a generic "stream dropped" (#578) (#608)
When WhisperX (or any ASR backend) failed to load its model, the transcribe SSE stream dead-ended on a generic "Transcribe stream dropped … Likely ASR backend failed to load" message with no actionable cause. Two root causes, both fixed: 1. WhisperX loads lazily inside transcribe(), so a load failure (faster-whisper weights, CTranslate2/cuDNN mismatch, torch-2.6 weights-only VAD regression) was buried in per-chunk errors and retried on every chunk. Added ASRBackend.ensure_loaded() (no-op default; WhisperX triggers its lazy loader) and call it in the transcribe pre-flight so the genuine cause surfaces once, up front, as a structured error event. 2. The pre-flight and audio-load error paths closed the SSE stream with a bare `error` and no terminal `done`, so the browser's native EventSource connection-drop could race and win against the structured error — discarding the real cause. Every terminal error now emits `done`, and the frontend latches the structured cause so a connection drop can't overwrite it with the generic message. Adds a fail-before/pass-after regression test driving the stream's async generator through the ASR-load-failure path; updates the existing #516 fake backend to the new ensure_loaded() contract; CHANGELOG ### Fixed entry. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8f2c4bbc5c |
fix(tts): NFC-normalize text + dense-script-aware chunking for long-form quality (#502/#505) (#587)
Two defensive fixes for non-Latin / long-form synthesis quality: #502 (Vietnamese clone distorted/unintelligible): the /generate text path never NFC-normalized its input, so pasted decomposed (NFD) Vietnamese — base letter + combining diacritic instead of the single composed codepoint — reached the tokenizer/model as two characters and rendered as garbled speech. Normalize the input text to NFC at the endpoint (no-op for already-composed text), mirroring what the duration estimator already does so the estimate and synthesis agree. #505 (long-form 5+ min degrades — repeated/skipped/mispronounced): the chunker split purely by character count (800), but CJK/kana/Hangul pack ~1 char = 1 syllable, so an 800-char chunk is ~4-5 minutes of audio in a single shot — past the model's reliable range, where it starts repeating/skipping. When a chunk is predominantly dense-script, cap it to max_chars/2.5 so each chunk's spoken length stays bounded; Latin/spaced text is unchanged. Dense-script detection is by code point (no literal CJK in source — no-literal-CJK gate stays clean). Tests: _dense_char_count, _effective_max_chars (shrink-when-dense, unchanged-for- Latin, disabled-passthrough, floor), and that a 400-CJK-char string now splits (was one chunk) while a Latin paragraph still doesn't. Note: #502's exact distortion still wants a user sample to fully confirm; this is the defensive NFC fix that's correct regardless. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7393ae80f5 |
feat(design): seed pin / re-roll for designed voices (#526) (#577)
Voice design rolled a brand-new random seed on every synth, so tweaking an attribute also re-rolled the whole base timbre — you could never iterate on the "same voice, slightly different". #526 asks for the seed to be shown with a "keep this seed" control. - Backend: `/generate` already accepted `seed` and echoed `X-Seed`, but left `used_seed=None` when nothing supplied one (non-deterministic, unreproducible, empty X-Seed). Now it materializes a concrete random seed when none resolves, so every take is reproducible and the real seed is always returned and stored — this also helps the clone/profile paths, not just design. - Frontend: new store slice (`designSeed`, `keepSeed`); the design synth reuses the pinned seed when "keep this seed" is on (via `pickDesignSeed`) and reads the authoritative seed back from `X-Seed`. Design tab gains a Seed field + "keep this seed" checkbox + "New seed" (re-roll) button. Test: `pickDesignSeed` (pin when kept+valid, re-roll otherwise, range guard). i18n keys added to en.json (other locales fall back; parity probe is advisory). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8ed76b40a6 |
fix(lang): propagate profile/request language into generation + longform (#533/#505/#502) (#565)
Non-English voices drifted to English/wrong-language because the request's or profile's language wasn't reaching the model: - #533: generate_speech() read instruct/ref_text/seed from a resolved profile row but never row['language'] (and collapsed Auto→None), so a German archetype previewed in German yet generated in English on the user's own call (and via Docker/API). Fall back to the profile's stored language when the request didn't pin one; an explicit non-Auto request language still wins. (Frontend already sets the dropdown on profile-select; this is the authoritative backend fix.) - #505 (B2): the audiobook/longform synth hardcoded language=None, so the engine re-autodetected per chunk and a non-English clone flipped language mid-render. Add _resolve_default_language (request → profile → autodetect) and thread the resolved language through _build_synth/_prepare_synth/_render_longform_sse, the three longform request models, the preview path, and the resume manifest. Genuine Auto/unset behavior is unchanged. - #502 (partial): the duration estimator weights combining marks (U+0300–036F) at 0.0, so NFD/decomposed text under-allocated frames → rushed audio. NFC- normalize text at the estimator entry — fixes the whole diacritic-script class (no-op for precomposed text). (The residual "distorted" core still needs the reporter's sample; tracked separately.) Tests (fail-before/pass-after): profile language reaches the engine (German→de; explicit/Auto override semantics); longform synth gets the resolved language (→ja), not None; NFD vs NFC duration parity (Korean Hangul diverges ~3x pre-fix). Full suite: tests/ 1740 passed, backend/tests/ 114 passed. Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a7ab148483 |
fix(asr): float16-unsupported GPUs fall back to int8 instead of "no segments" (#561)
#551: both CTranslate2 ASR backends request compute_type="float16" on CUDA with NO fallback. On GPUs without efficient fp16 (older Maxwell/Pascal, GTX 16xx) or a CTranslate2/cuDNN binary mismatch, WhisperModel/whisperx.load_model raise a ValueError at construction — which escaped the existing OOM-only `except RuntimeError`, so every chunk failed and the user got "Transcription produced no segments". Add a per-device compute_type fallback chain (cuda: float16 → int8_float16 → int8; cpu: int8 → float32) to both backends + the ASR sidecar, alongside (not replacing) the existing OOM→CPU path, with an ASR_COMPUTE_TYPE override for exotic hardware (documented in README). Also in the same ASR-robustness pass: - #549: PyTorchWhisperBackend._ensure_pipe wraps the transformers pipeline load and re-raises an actionable error (reinstall transformers / use faster-whisper) instead of a bare "Could not import module 'AutoFeatureExtractor'". - #516: the /dub/transcribe SSE generator is wrapped so it can NEVER close without a terminal event — any unanticipated exception now yields a structured `error` (with build_failure's hint) + `done`, turning "stream dropped, likely ASR failed" into the real cause + Retry. - failure.py: COMPUTE_TYPE_UNSUPPORTED + TRANSFORMERS_IMPORT classes so the no-segments toast is actionable. Tests (fail-before/pass-after): float16-unsupported → int8 for both WhisperX + FasterWhisper; a generic non-OOM RuntimeError still raises; classify() maps the two new classes; the SSE stream always terminates with error→done. 7 + 1 passed, 17 in the failure suite (no regression). Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
3656f0a4ef |
fix(profiles): decouple design-profile save from TTS render (#476) (#488)
* fix(profiles): decouple design-profile save from TTS render (#476) Saving a design voice profile forced a full TTS model load + inference to render a deterministic identity sample. On a fresh model-less image (Docker first-run) that 503'd, so the save failed. A secondary guard also rejected an all-Auto design (empty instruct) with a 422. Saving a design profile is now a pure persistence operation: - The seed-42 identity sample render is attempted opportunistically but is non-fatal — if the engine isn't ready the row is persisted with ref_audio_path=NULL (sample pending). The row's vd_states + instruct already make the voice fully usable (generation.py falls back to instruct-only conditioning for design profiles with no ref audio). - The sample is rendered lazily + cached on the first GET /profiles/{id}/audio request; if the engine is still unavailable that path returns a precise "model not ready — finish setup / download a model" 503. - The all-Auto (empty-instruct) design is now saveable (vd_states still required). Adds tests/test_profile_design_save_decouple.py (top-level tests/, asyncio.run per test) covering: design save with model unavailable creates the row instead of 503-ing; all-Auto design is saveable; the pending sample materializes on first /audio request. Updates the unification spec (docs-sync). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): contain profile-audio paths under VOICES_DIR (CodeQL CWE-22) The lazy design-sample path was built as `os.path.join(VOICES_DIR, f"{profile_id}.wav")` / `os.path.join(VOICES_DIR, audio_file)` where profile_id is the request path param — CodeQL flagged 5 high-severity path-injection alerts (profiles.py + the taint flowing into archetypes.py's torchaudio save). Add `_safe_voice_path()` (basename + safe-char sanitise + realpath containment, mirroring core.config.dub_seg_path) and route both the read and lazy-render sites through it; a traversal id now 404s instead of escaping VOICES_DIR. Regression test covers the containment guard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): use CodeQL-recognized path-injection guards (CWE-22) The previous `_safe_voice_path()` helper was correct (basename + realpath containment) but CodeQL's taint tracking didn't propagate the barrier through the function return, so the 5 path-injection alerts persisted. Switch to guards CodeQL recognizes, inline at each file-op site: - validate `profile_id` against the generated-id charset (`[A-Za-z0-9_-]{1,64}`) with `re.fullmatch` and 404 on mismatch (covers the `f"{profile_id}.wav"` render path); - read only `os.path.join(VOICES_DIR, os.path.basename(name))` so a stored/derived filename is always a direct child of VOICES_DIR (covers the read + the taint flowing into archetypes.py's torchaudio save). Drop the helper. Test now asserts a traversal/separator/NUL profile_id 404s at the guard. Same security property, recognized by CodeQL. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): inline realpath+commonpath containment for CodeQL (CWE-22) CodeQL didn't recognize the earlier sanitizers — neither the helper (barrier hidden behind a function return) nor os.path.basename / a cross-function regex guard cleared the 5 path-injection alerts. Use the canonical, CodeQL-recognized form INLINE at each file-op site: resolve the path with os.path.realpath (which collapses any `..`) and confirm os.path.commonpath((base, path)) == base before the read / the render, returning 404 / raising on escape. Same property the helper had, now in a shape CodeQL's taint tracking follows. Keeps the profile_id charset guard as defense-in-depth. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(profiles): route design-sample path through shared _voices_path guard (#476) The inline realpath+commonpath containment in get_profile_audio and _materialize_design_sample wasn't recognized by CodeQL as a path-injection sanitizer (5 new high-severity py/path-injection alerts at the file-op sites, incl. archetypes.py mkdir via the rendered Path). Both now reuse the existing _voices_path() helper, which applies the os.path.basename() barrier plus symlink-resolved containment — the same guard the consent endpoint uses and that CodeQL already accepts. Behavior is unchanged: the DB columns only ever hold bare {profile_id}.wav filenames, so basename() is a no-op here. Tests: tests/test_profile_design_save_decouple, test_profile_unification, test_profile_consent, test_archetype_blank_guard — 25 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <test@local> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1650a121db |
fix(dub): auto-assign per-speaker voices in multi-speaker dubbing (#486) (#490)
Multi-speaker dubs detected speakers and built per-speaker clones (Voice
dropdown showed "From Video → Speaker N"), but most segments stayed on
"Default" voice and had to be set by hand — inconsistently across runs.
Root cause: after diarization, dub_core stamped each long line (the
default-on per-segment-ref path) with `auto-seg:{id}` as its profile_id.
The dub editor's Voice <select> (and the Cast panel) only render `auto:`
options, so an `auto-seg:` value matched no <option> and silently showed
"Default". Short lines (<3s) fell through to `auto:{speaker}`, which DID
render — hence "sometimes the cloned voice is picked".
Fix: bind every segment to the UI-visible `auto:{speaker}` whenever its
detected speaker has a clone; only fall back to `auto-seg:{id}` when the
speaker has no per-speaker clone at all. The per-segment-ref quality win
is preserved: dub_generate's `auto:` branch now transparently prefers
THIS segment's own per-segment ref (segment_clones[seg_id]) when present,
else the per-speaker clone. Manual overrides and the no-clone path are
untouched; existing jobs that persisted `auto-seg:` ids still resolve.
Tests: tests/test_dub_multispeaker_voice_486.py — assignment binds to
auto:{speaker} (not auto-seg:), never clobbers manual overrides, falls
back to auto-seg: only when the speaker has no clone; generate-time
resolution prefers per-segment ref then per-speaker clone. Green
alongside the existing dub generate/incremental/segmentation suites.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
4e3136c1e0 |
fix(translate): guess source language from text instead of defaulting to "en" (#478)
When neither the request nor the job carries a detected source language, _resolve_source_lang() silently fell back to "en". For non-English audio (e.g. Korean) this produced en -> en, which has no Argos package and failed every segment — even though WhisperX had detected the language correctly (e.g. "Detected language: ko (0.98)"). Add a last-resort script-based guess (ko/ja/zh/ru/ar) from the segment text so the bare "en" fallback no longer breaks non-English dubbing. Co-authored-by: stronghamjji <289942360+stronghamjji@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
18793a99a7 |
feat(audiobook): durable crash-resume for interrupted longform renders (#470)
* feat(audiobook): durable crash-resume for interrupted longform renders
Chapter WAVs were already content-addressed (a re-run reused finished chapters),
but resume only worked if the user could re-submit the EXACT script — impossible
for Stories, whose plan is compiled from cast+lines. This persists the plan
itself so an interrupted render is resumable without the original input.
- New services/longform_resume.py (pure file/JSON): on render start, write a
resume.json manifest (compiled plan + render params + title) into the job work
dir, atomically; clear it on successful completion. read/has/clear/build
helpers, schema-versioned (a foreign/corrupt manifest is ignored, never
resumed).
- _render_longform_sse: accepts an optional job_id + resume flag (resume reuses
the original job row + cached chapters instead of creating a new one); writes
the manifest at start, clears it on done. Both front doors (/audiobook,
/longform/render) unchanged for callers.
- GET /audiobook/jobs — lists interrupted renders (running/failed longform jobs
that still have a manifest; a job left "running" across an app restart is
interrupted by definition), with title + total/done chapter counts for the UI.
- POST /audiobook/resume/{job_id} — rebuilds the plan from the manifest and
replays _render_longform_sse under the original job_id; the content-addressed
cache makes finished chapters instant, so only the unrendered ones synthesize.
404 on unknown id / missing manifest.
Resume durability is best-effort — a manifest failure never blocks the render.
The resume UI affordance is a follow-up (the endpoints are ready for it).
Tests: tests/test_longform_resume.py (7, pure manifest round-trip / version &
corrupt rejection / atomic write — monkeypatches OUTPUTS_DIR, no global
core.config stub so the shared tests/ session isn't polluted) +
backend/tests/test_audiobook_resume_api.py (6, config-stub: jobs-list with
progress, failed-included, done/manifestless/non-longform excluded, resume
404s). 13 passed. CJK green. Stale module docstring updated.
* fix(audiobook): confine resume paths — py/path-injection (CodeQL) + quality
The default-setup CodeQL (security-and-quality suite) flagged the crash-resume
work: longform_resume built filesystem paths from job_id, which on the
POST /audiobook/resume/{job_id} endpoint is a request-supplied path param →
py/path-injection (10 high-severity sinks: open/replace/remove/makedirs/isfile).
- longform_resume.work_dir now confines like profiles._voices_path: reject an
unknown job_type or an id that isn't a bare safe token (^[A-Za-z0-9_-]{1,64}$),
then realpath + startswith(OUTPUTS_DIR + os.sep) — a crafted id (`../`, NUL,
separators) can never escape OUTPUTS_DIR. Returns None on violation; all
callers (manifest_path/read/write/clear/has) degrade gracefully.
- The resume endpoint also gates the path-param id up front (404 on a bad
token) — barrier at the source as well as the sink.
Also cleared the quality alerts the same diff introduced:
- py/repeated-import: the 4 inline `from services import longform_resume` calls
collapse to one module-top import (it's pure, no torch).
- py/empty-except: the best-effort manifest blocks now logger.debug instead of
a bare `pass`.
13 resume tests still pass; all job ids in tests are safe tokens.
* fix(audiobook): sanitize resume job_id at the source (path + log injection)
The first CodeQL pass wasn't enough: resume made job_id request-controlled, so
it tainted not just the manifest paths but the EXISTING work-dir join and the
progress log lines too (py/path-injection + py/log-injection, ~14 alerts).
Fix at the source so the whole dataflow is clean:
- _render_longform_sse strips job_id to a safe token (`re.sub` removing anything
but [A-Za-z0-9_-], capped 64) right after it's resolved — no path separator,
no CR/LF can survive, whether the id came from the resume path param or a
fresh uuid.
- The work dir now routes through longform_resume.work_dir, which adds the
proven os.path.basename(seg)==seg barrier (the shape CodeQL accepts in
_voices_path) on top of the realpath+startswith confinement — so the join and
every path derived from it (meta/concat/out) is sanitized.
- The best-effort manifest-write log no longer interpolates the raw exception
(uses exc_info); clear_manifest's OSError handler returns instead of bare pass
(py/empty-except).
13 resume tests still pass.
* fix(audiobook): launder resume job_id via trusted FS scan (CodeQL path/log-injection)
The custom realpath/regex barriers weren't in CodeQL's recognized sanitizer set,
so the request-supplied resume job_id kept tainting the work-dir/manifest paths
and the progress logs. Switch to the pattern CodeQL does accept — launder the id
through a trusted filesystem enumeration:
- longform_resume.scan_resumable() lists resumable jobs by scanning OUTPUTS_DIR
for <type>_<id>/resume.json; every id it returns is sourced from os.listdir
(never request input).
- POST /audiobook/resume/{job_id} now only resumes an id that scan_resumable()
reports (membership match), and uses the (job_type, job_id) pair FROM that
trusted list for everything downstream — so nothing request-controlled reaches
a filesystem path or a log line.
- GET /audiobook/jobs lists from scan_resumable() too (filesystem-sourced ids).
work_dir keeps the realpath+startswith+basename confinement as genuine defense;
the render path's job_id is now always either a fresh uuid or a laundered id.
13 resume tests still pass.
* fix(audiobook): exact-match allowlist on the work-dir name (CodeQL path-injection)
The remaining 4 path-injection alerts were inside work_dir: I validated job_id
with an anchored regex but then joined a DIFFERENT f-string (`{job_type}_{job_id}`),
so CodeQL didn't carry the sanitization to the joined value. Mirror the pattern
the repo's _safe_cover_path uses (which CodeQL accepts): validate the WHOLE
joined component against an exact-match allowlist regex (_SAFE_SEG_RE), then
confine with os.path.commonpath containment (the recognized barrier) instead of
startswith. 13 resume tests still pass.
* fix(audiobook): basename-sanitize the work-dir name for CodeQL path-injection
The exact-match regex alone wasn't credited; route the joined value through os.path.basename() first — the sanitizer CodeQL recognizes (mirrors _safe_cover_path) — then the regex + commonpath. Functionally identical (no separator in the name) but clears the 4 remaining alerts. 13 tests pass.
* fix(audiobook): allow-list membership guard launders resume job_id (CodeQL)
The next(... if pair[1]==job_id) comparison-select didn't sanitize for CodeQL. Build a dict of resumable ids from the trusted scan and gate with 'if job_id not in resumable' — the membership barrier CodeQL recognizes — then use job_id directly downstream. 13 tests pass.
* fix(audiobook): eliminate request→path flow in resume (definitive CodeQL fix)
Five rounds of recognized path-injection barriers (regex, basename, exact-match,
commonpath, membership-guard) still left CodeQL flagging the resume job_id →
work-dir/manifest/log flow. Remove the flow entirely instead of guarding it:
- scan_resumable() now returns {job_type, job_id, manifest_path} where
manifest_path is built from the os.listdir dir name (trusted), plus
load_manifest_file(path) / discard_manifest_file(path) that operate on those
trusted paths. The request job_id is used ONLY to *select* a scan entry, never
to build a path.
- POST /audiobook/resume/{job_id} reads the manifest via the trusted scan path
and renders under a FRESH server uuid (job_id=None). The chapter cache is
content-addressed (keyed by chapter content, not the job id), so finished
chapters still hit instantly — resume works, but the request's id never names
a work dir, output file, or log line.
- The interrupted job's manifest is discarded (trusted path) once the fresh-id
resume kicks off, so it stops showing as resumable.
Net: no request-controlled value reaches any file operation or log on the
render path (job_id there is always a server uuid). work_dir keeps its
confinement barriers as defence-in-depth. 13 resume tests pass.
|
||
|
|
2b8c8aec7c |
fix: actionable errors for non-executable engine binary (#437) + unreachable backend (#438/#454/#466) (#471)
Two reliability bugs from open issues, both first-run papercuts where the error told the user the wrong thing. #437 — `[Errno 13] Permission denied: bin/omnivoice-tts-linux-x86_64`: a git clone / zip extract on POSIX can drop the bundled binary's execute bit. It only surfaced at spawn time, and the generic synth handler then mislabeled it as "ran out of memory" and told the user to flush the model. - omnivoice_gguf.is_available() now self-heals: after the SHA check confirms the binary is the right file, it adds +x (best-effort) on POSIX; if it can't, it returns a clear "isn't executable — run chmod +x <path>" message instead of a spawn-time crash. No-op on Windows. - generation.py classifies PermissionError / EACCES / "Permission denied" as its own case ("a bundled binary lost its execute bit — reinstall or chmod +x"), so it never again masquerades as OOM. #438/#454/#466 — bare "Failed to fetch" / "NetworkError": when the local backend is still starting, crashed, or the dev server dropped, fetch() throws a TypeError that propagated raw to the user. - client.ts apiFetch now catches the thrown fetch and raises an ApiError with an actionable message ("Can't reach the local OmniVoice backend — it may still be starting up… restart the app or check Settings → Logs"), status:0 to mark a transport failure vs an HTTP error. Tests: client.test.ts +1 (thrown fetch → ApiError status 0 + actionable text); 3 pass. CJK guard green. |
||
|
|
35c063ae52 |
feat(persona): /personas export·import·inspect router + wiring (#29 slice B) (#461)
* feat(persona): .ovsvoice build/parse core + embed_watermark(force=) (#29 slice A) Extends the merged persona-bundle nucleus (constants, normalize_spdx, build_manifest, build_consent_json) with the model-coupled core that the export/import router (next slice) will sit on: - `build_persona_bundle(profile, *, license_spdx, tags, include_reference, embed_fn, …)` → assembles the .ovsvoice ZIP in memory: a watermarked preview.wav (24 kHz mono 16-bit, downmixed + resampled + trimmed ≤8 s), manifest.json, a legacy-shaped metadata.json (so an older OmniVoice can still import the ref audio), optional consent.json, and the raw ref/locked/consent members unless include_reference=False (privacy / preview-only, A12). Raises NoPreviewSource (router → 503) when no source clip is readable (A2-A5). - `parse_persona_bundle(bytes)` → validates the ZIP, prefers manifest.json and falls back to legacy metadata.json, resolves audio members by prefix (last-wins, B9; member names never build paths — zip-slip safe), normalizes the SPDX id, flags preview-only / future-schema_version. Raises BundleError(400|413) for B1-B11. No DB, no file writes. - `ParsedPersona` dataclass with `extract_member(prefix, dest_path)` — the router derives dest_path from the server-generated id, never the member name. - `embed_watermark(..., *, force=False)`: keyword-only flag that bypasses the user's invisible-watermark preference for the mandatory persona preview, but still no-ops without AudioSeal. All existing positional call sites are unchanged (default force=False) — default cross-platform behaviour identical. All heavy imports (torch/torchaudio/watermark/audio_io) are lazy so the module stays model-free at collection (avoids the local torch/Triton segfault). tests/test_persona_bundle.py: +31 cases — parse validation (manifest/legacy selection, preview-only, future-schema, missing/malformed/no-audio → 400, oversize → 413, bad-SPDX normalize, last-wins dup, advisory consent), build round-trip (identity fields, metadata sibling, no-source → NoPreviewSource, include_reference=False, stereo/off-rate downmix+resample), and the force= unit (D1/D3). 25 pure cases pass locally; the 6 torchaudio-coupled cases run on CI (local torch+pytest segfault is pre-existing). CJK guard green. * feat(persona): /personas export·import·inspect router + wiring (#29 slice B) Thin HTTP layer over the persona_bundle service (slice A), registered in main.py next to the legacy marketplace router: - POST /personas/export/{id} → builds the .ovsvoice off the event loop (run_in_executor) and streams it (application/zip, .ovsvoice filename; empty name → persona_<id>). 404 when the profile is missing; NoPreviewSource → 503 (no readable source audio); any other build error → 503 with a generic message (no raw exception text in the body). - POST /personas/import → parse (BundleError → its HTTP status), extract audio members to server-named files ({id}{ext}/{id}_locked{ext}/{id}_consent{ext} — never the member name, zip-slip safe via profiles._voices_path), 17-column INSERT (legacy 13 + the 4 consent columns), event_bus emit after commit. Verified-own-voice is granted ONLY with a real recording ≥ floor AND non-empty consent_text AND consent.json present (forgery guard, B12-B16). Rollback: every written file is deleted on any extraction/INSERT failure; id-collision retries once (renaming the on-disk files to the new id). Accepts legacy .omnivoice too (case-insensitive extension guard). - POST /personas/inspect → manifest + consent summary with NO DB row and NO file extracted (import-preview UI). backend/tests/test_personas_api.py: 13 cases (config-stub pattern → mounts only the router, no main/torch import) — export 404; import bad-ext/non-zip/missing- manifest 400; round-trip row+file under server name; case-insensitive ext; forgery-unverified; verified-with-recording; short-recording-unverified; preview-only-as-ref; legacy .omnivoice; inspect no-write + consent summary. 13 passed locally. CJK guard green. |
||
|
|
ca8a2e8eb8 |
feat(audiobook): PDF ingest for /audiobook/import (ebook-in core value) (#459)
The audiobook importer accepted .txt/.md/.epub but not PDF — the single most common "ebook in" format. Add a pure `pdf_to_chapter_script(data)` that extracts the text layer page-by-page and runs it through the existing chapterizer, so PDFs land in the same `# Heading` + body grammar EPUB and plaintext already produce (one front door onto the unchanged render pipeline). - Dep: `pypdf>=4.0` — pure-Python, MIT, zero native deps, so PDF import behaves identically on macOS/Windows/Linux (default-feature cross-platform rule). EPUB + plaintext stay stdlib-only; only PDF needs a real parser. - Robustness, surfaced as actionable 400s rather than silent empty imports: corrupt file, password-protected (empty-password decrypt attempted first), scanned/image-only (no text layer → clear "scanned PDF" message), and a page-count ceiling. A single unparseable page is skipped, not fatal. - Route: `.pdf` branch in audiobook_import; frontend accept filter + api-client doc updated to `.txt,.md,.epub,.pdf`. tests/test_longform_import.py: 5 PDF cases (extract+chapterize, no-marker single chapter, corrupt, image-only, page-cap) using a hand-built in-memory PDF — no PDF-authoring test dep, mirroring the in-memory-EPUB approach. 16 passed; frontend suite 401; CJK guard green. |
||
|
|
4531e999b1 |
feat(capture): opt-in LLM refinement on REST /transcribe (parity with live dictation) (#457)
The live-dictation socket (capture_ws) already runs the final transcript through the configured local LLM (disfluency/self-correction/punctuation cleanup, Wave 2.1). The REST /transcribe endpoint — the MCP / CLI / file-upload surface — only did the always-on hallucination-loop collapse, so agentic and batch callers couldn't get the same cleaned output. Add an opt-in `refine` form flag that runs the identical `maybe_refine` pipeline off-thread: - OFF by default → existing MCP/CLI callers keep raw-only output and pay no LLM latency (backward-compatible). - Honours the user's Settings → Dictation-refinement config and silently passes through when no LLM backend is configured (cross-platform default parity — identical no-op everywhere with no LLM). - Raw `text` is always returned; `refined_text` is added only when the LLM actually changed the text — same contract the socket emits. tests/test_capture_refine.py: 13 cases — flag-off no-call, refined_text on change, no-op/identical omission, and flag parsing. maybe_refine is patched at its source module since the handler imports it lazily. |