Only the TTS model (~2.4 GB) is required on first run; ASR models are
per-platform curated picks (curated_on in models.yaml) installed on demand.
Every transcription surface returns a typed asr_model_missing error with a
one-click download CTA instead of silently pulling multi-GB Whisper weights.
Settings -> Models is a grouped, platform-aware catalog. New guided
permissions UX (wizard System Check + Settings -> Permissions + mic
pre-flight) with native mic-state checks and OS settings deep-links. New
parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated
capture preference so multilingual dictation never regresses). Docs:
expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ.
Hardening: preflight fails open for custom model pins, ROCm curation no
longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged
consent key, CaptureWidget setup race fixed, offline-cache CI simulation
fixes so empty-cache runners stay green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Residual A — the chunked dub-stream had a PARALLEL wedge mechanism (its own
ping-loop timeout, its own _reset_pool_on_wedge, a dead-end "Try restarting
the server" message). A wedged chunk now routes through the SAME
run_transcribe_guarded bound+reset as the whole-file paths (#731/#851): the
guard resets the poisoned pool once per wedged attempt (no double-reset on
retry) and the user sees the actionable ASRTimeoutError. The reset logic is
extracted to asr_backend.reset_pool_after_wedge — one shared mechanism, so
the semantics can't drift again. run_transcribe_guarded also gains a
timeout_env param so chunk errors name OMNIVOICE_TRANSCRIBE_CHUNK_TIMEOUT_S
instead of the whole-file knob.
Residual B — the crash-isolated ASR sidecar (#393, faster-whisper-isolated)
is wired as an explicit ESCAPE HATCH, not a default:
- selectable end-to-end: Settings engine list gets an explanatory
install_hint; honest gpu_compat ("cuda","cpu" — it wraps the same
CTranslate2 engine as faster-whisper); get_active_asr_backend now hands
back a process-wide singleton for subprocess-isolated backends (a fresh
instance per request would leak atexit hooks and respawn the sidecar —
reloading its model — on every transcribe).
- on the SECOND consecutive guarded timeout-with-reset in one session
(resets aren't recovering the hang; the wedged thread keeps its VRAM),
the error the user sees + the log recommend switching to the isolated
engine in Settings → Engines. Never auto-switched (owner rule: no silent
behavior divergence); a completed transcribe resets the streak.
Tests (fail-before/pass-after verified against origin/main): wedged-chunk
SSE integration (reset count + actionable error + recommendation surfaces),
consecutive-timeout streak (fires at 2, resets on success, suppressed when
already on the isolated engine), timeout_env parametrization, shared-reset
helper, isolated backend in list_backends with hint + honest availability,
singleton caching, gpu_compat matrix entry. Docs: troubleshooting §14 gains
the chunk knob + escape-hatch guidance.
Closes the residuals tracked on #730.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Builds on the device probe from PR 1. Backend-only; the routing keys are
wired into /engines in PR 3.
- #390 closed: MLXAudioBackend / MLXWhisperBackend now call the shared
`core.device_caps.mlx_supported()` gate FIRST, before importing the
package. On Linux/Windows/mac-Intel they report unavailable and never
advertise a usable `mps` route, even with a stray mlx wheel installed.
Replaces the ASR backend's ad-hoc inline MPS check with the one shared
rule. (The Wave-4.4 OSError/RuntimeError import-guard is preserved — it
now lives behind the platform gate; its test forces the gate open so the
guard stays the path under test.)
- `ASRBackend` ABC gains `gpu_compat: tuple[str, ...] = ("cpu",)` mirroring
TTSBackend, and each subclass declares its real targets:
whisperx/faster-whisper → (cuda,cpu); mlx-whisper → (mps,cpu);
pytorch-whisper → (cuda,mps,cpu); nemo/funasr → (cuda,cpu);
moonshine → (cpu,). Inert until PR 3 serializes them.
- IndexTTS2 declares `gpu_compat = ("cuda","cpu")` so it stops advertising
the inherited CPU-only default.
- ROCm is deliberately NOT claimed for any ASR engine (or for IndexTTS2):
CTranslate2 has no upstream HIP build, and an unverified `rocm` claim
would route ROCm hosts to a broken GPU path — strictly worse than the
honest `cpu_fallback` the resolver already emits ("declares CUDA only;
ROCm not in its compat set"). The per-engine TTS ROCm audit is a tracked
follow-up that will verify each path before claiming it.
Tests: MLX gate regression (both backends, on/off Apple), ASR gpu_compat
tuples + no-false-rocm invariant, IndexTTS2 override; existing MLX
import-guard test updated for the new gate ordering.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>