Only the TTS model (~2.4 GB) is required on first run; ASR models are
per-platform curated picks (curated_on in models.yaml) installed on demand.
Every transcription surface returns a typed asr_model_missing error with a
one-click download CTA instead of silently pulling multi-GB Whisper weights.
Settings -> Models is a grouped, platform-aware catalog. New guided
permissions UX (wizard System Check + Settings -> Permissions + mic
pre-flight) with native mic-state checks and OS settings deep-links. New
parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated
capture preference so multilingual dictation never regresses). Docs:
expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ.
Hardening: preflight fails open for custom model pins, ROCm curation no
longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged
consent key, CaptureWidget setup race fixed, offline-cache CI simulation
fixes so empty-cache runners stay green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three-part fix for the v0.3.5-comparison report:
1. Perf: since v0.3.6 (#308), a clone reference without a stored
transcript triggered a FULL ASR model load + transcribe on every
/generate — get_active_asr_backend() builds a fresh whisper backend
per call. Measured live: 92.7s wall vs 14.9s of actual TTS. Now the
first auto-transcript is persisted onto the (unlocked, clone-kind)
profile row, and transcribe_reference caches results by audio
content hash (bounded LRU, no model/VRAM held), so the cost is paid
once per clip, not per request. User-typed transcripts are never
overwritten; locked/design profiles are excluded from the persist.
2. Clear History: the workspace UX overhaul (#374) moved history into
the right-side WorkspaceHistory panels and dropped the old Sidebar's
clear-all control (the Sidebar is now hidden in every mode). Both
the Voice and Dub panels get a scoped Clear History button wired to
the existing DELETE /history and /dub/history endpoints, with the
same confirm dialog the Sidebar used.
3. Auto-play: the finished-render playback (playBlobAudio) has no
on-screen player and the only stop lived in the Voice ActionBar's
CTA morph — unstoppable from the Dub workspace, profile pages, or
after navigating away. A global PlaybackStopPill now appears for any
'output' playback on every page. The existing Settings → Appearance
"Auto-play preview" pref (#667) now also gates the generate path,
as its label always promised (default ON — no behavior change).
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Voice cloning without a transcript fell through to OmniVoice's built-in
load_asr_model() — a transformers pipeline() load of
whisper-large-v3-turbo that fails outright on transformers 5.3 — even
when whisperx / faster-whisper / mlx-whisper were installed and working.
The dub pipeline already used the registry; the /generate clone path
never did.
- services/asr_backend.py: new transcribe_reference() resolves the
active registry backend (honoring auto-detect order and the
OMNIVOICE_ASR_BACKEND override), extracts text from either result
shape (top-level "text" or whisperx-style segments), and degrades to
None on any failure so the model fallback behaves exactly as before.
When the registry itself resolves to pytorch-whisper it defers to the
model's lazy load instead of building a second pipeline.
- api/routers/generation.py: transcript-less references get transcribed
in the GPU pool before inference.
- tests/test_transcribe_reference.py: covers both result shapes,
failure degradation, and the pytorch-whisper deferral.
The remaining half of #308 — pytorch-whisper itself being incompatible
with transformers 5.3 when it truly is the last resort — is tracked in
the issue.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>