Files
VoiceStudio/docs/engines
Palash DebnathandClaude Fable 5 2d5f2e800e feat(omnivoice): voice prompts that survive restarts + opt-in FlashInfer (~2.2x) (#1565)
* feat(omnivoice): port upstream VoiceClonePrompt persistence + FlashInfer opt-in

Upstream k2-fsa teardown ports, verified with generated voice samples:

- VoiceClonePrompt.save()/.load() (upstream format v1, weights_only-safe)
  on the vendored model, and a disk layer under the in-memory prompt LRU
  (DATA_DIR/prompt_cache, keyed by ref path+mtime+ref_text+preprocess,
  32 newest kept, OMNIVOICE_PROMPT_DISK_CACHE=0 opts out). First generation
  of a session with a known voice skips the reference re-encode and any
  auto-transcription pass — verified across two real processes (encodes=1
  then encodes=0, same voice).
- omnivoice_flashinfer.py ported (packed CFG attention, fused kernels,
  optional CUDA graphs), schedule adapted to our num_step+1 divergence.
  Opt-in via OMNIVOICE_FLASHINFER=1|graph, CUDA-only, replaces
  torch.compile for the session; missing package / apply failure / runtime
  failure all degrade with a named reason (same #278 contract as compile:
  classify → unapply → retry once, session latch). Measured 2.20x at
  batch=1 on an RTX 4090 with byte-identical text and clean ASR round-trip.
- Docs: OmniVoice guide gains instruct+reference combination semantics
  (consistent instruct stabilizes cloning, reference wins conflicts),
  inline pronunciation control (pinyin / CMU), prompt persistence, and
  corrects the 'no voice design' claim; performance.md documents both new
  env knobs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: point changelog entries at the real PR number (#1565)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pr): harden FlashInfer lifecycle + prompt-cache writes per review

Bot harvest round 1 (#1565): unapply on apply-failure (half-patched model
could crash the next render); pin eager-mode FlashInfer inference to one
thread too — the attention plan and packed position ids are per-generation
module state, so interleaved _gpu_pool workers would corrupt each other;
restore the CAPTURED pre-apply attention impl (could be flash_attention_2)
instead of assuming sdpa; unique tmp name per prompt-cache write; correct
the _forward_logits layout docstring; resolve VoiceClonePrompt at test
runtime; docs — Known limits keeps only the limitation, performance.md
states the VRAM cost and scopes the fallback claim to classified kernel
failures.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pr): round-2 review — publish only a fully restored model, redact latch reason, tighten CPU-persistence test

Greptile: the runtime fallback now unapplies BEFORE swapping generate, so
a concurrent render keeps queuing behind the thread-affinity wrapper while
teardown mutates modules. CodeRabbit: FlashInfer failure reasons pass
through core.failure.sanitize before latching/logging (wheel paths embed
the user's home); the save-portability test now creates the tokens on CUDA
when available and asserts the persisted payload itself is CPU-resident.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(pr): fail-closed latch reason when the sanitizer itself breaks

CodeQL empty-except + CodeRabbit round 3: if core.failure.sanitize raises,
the raw reason (home paths, wheel paths) was latched anyway. Now only the
exception class survives with a fixed redaction note; two regression tests
(normal redaction + sanitizer failure).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 16:25:29 +00:00
..

Engine guides

One page per engine: what it's for, what it needs, how to enable it, and its quirks. Select engines in Model Catalogue → Engines (or quick-switch with Ctrl/Cmd+E), or pin one with OMNIVOICE_TTS_BACKEND / OMNIVOICE_ASR_BACKEND.

The compute device (CUDA/ROCm/MPS/CPU) is auto-detected; pin it under Settings → Performance & Device (or OMNIVOICE_DEVICE) if auto-detect picks wrong — see performance.

Measured speed/VRAM numbers live in benchmarks; what each engine can do expressively in expressive-speech; sidecar disk footprints in disk-usage; the bar a new engine must clear in engine-acceptance.

New to VoiceStudio? Install the app first — macOS (first launch needs the one-time right-click → Open Gatekeeper approval), Windows, Linux, Docker.

Text-to-speech

Engine Guide Runs on Cloning Enabled by
VoiceStudio (OmniVoice) — default omnivoice CUDA · MPS · CPU ✅ installed by default
VoxCPM2 voxcpm2 CUDA · MPS · CPU ✅ + voice design pip install "voxcpm>=2.0.3"
MOSS-TTS-Nano moss-tts-nano CUDA · CPU ✅ (ref only) clone + uv pip install -e .
KittenTTS kittentts CPU — (8 preset voices) pip install kittentts
MLX-Audio (Kokoro, CSM, Dia, …) mlx-audio Apple Silicon model-dependent pip install mlx-audio
CosyVoice 3 cosyvoice CUDA · CPU ✅ clone + requirements
GPT-SoVITS gpt-sovits external server ✅ its own API server
Sherpa-ONNX sherpa-onnx CUDA · CPU — pip install sherpa-onnx + model dir
IndexTTS 2.5 indextts CUDA · CPU ✅ + emotion one-click sidecar install
OmniVoice GGUF omnivoice-gguf CUDA · MPS · CPU ✅ bundled binary
Supertonic-3 supertonic3 CPU — (7 preset voices) uv sync --extra supertonic + license
MOSS-TTS-v1.5 (8B) moss-tts-v15 CUDA · CPU ✅ clone + env var
dots.tts (2B) dots-tts CUDA · CPU (not Windows) ✅ clone + env var
OmniVoice (subprocess) omnivoice-subprocess CUDA · MPS · CPU ✅ opt-in pick, no install
PocketTTS (Kyutai) pockettts CPU (not Intel Mac) ✅ uv sync --extra pockettts + license
Confucius4-TTS confucius4-tts CUDA · CPU ✅ clone + env var

Speech-to-text

Engine Guide Runs on Best at Enabled by
WhisperX whisperx CUDA · CPU dubbing (word timestamps + diarization) installed by default
Faster-Whisper faster-whisper CUDA · CPU general transcription installed by default
Faster-Whisper (isolated) faster-whisper-isolated CUDA · CPU unattended batches opt-in pick
MLX Whisper mlx-whisper Apple Silicon Mac default pip install mlx-whisper
PyTorch Whisper pytorch-whisper CUDA · MPS · CPU ROCm hosts installed by default
Parakeet TDT (NeMo) nemo-parakeet CUDA · CPU 25 languages, fast CPU separate venv (never the app's)
Parakeet TDT (MLX) parakeet-mlx Apple Silicon dictation, 25 EU languages default on mac-ARM source installs
Moonshine moonshine CPU edge/low-power, no timestamps pip install (see guide)
FunASR (SenseVoice) funasr CUDA · CPU 50+ languages, inline diarization pip install funasr
Sherpa-ONNX dictation sherpa-onnx-asr CPU live streaming dictation curated model download
OpenAI-compatible (remote) openai-compatible-asr network offloading to a server (audio leaves the machine) Model Catalogue

Speaker diarization is not an engine registry of its own — the dub pipeline uses pyannote (HF-gated; see diarization) and FunASR can diarize inline with its cam++ speaker model.