Files
VoiceStudio/docs/engines
Palash Debnath 180371910e fix(cosyvoice): patched dependencies, no wetext download, Windows long-path hint
Built the real environment on Windows to check the one-click install.
- The pins upstream carries with published advisories (diffusers,
  hydra-core, lightning, modelscope, onnx, protobuf, transformers) are
  raised to fixed releases; that exact set installs and passes the
  installer's import probe. pyarrow, not needed, is dropped. A test keeps
  every pin at or above its advisory fix.
- wetext is dropped with its install-time fetch. Its data exists only on
  ModelScope, which rate-limits downloads; a throttled fetch left files
  missing while reporting success, so the normaliser failed silently.
  CosyVoice reads text as written without it, and nothing reaches
  ModelScope at install or synthesis time. The unused post-install hook
  goes with it.
- openai-whisper builds from source, and under a long cache path the build
  hits Windows' 260-character limit; the install failure now says to turn
  on long-path support.
2026-09-10 11:39:50 -07:00
..

Engine guides

One page per engine: what it's for, what it needs, how to enable it, and its quirks. Select engines in Model Catalogue → Engines (or quick-switch with Ctrl/Cmd+E), or pin one with OMNIVOICE_TTS_BACKEND / OMNIVOICE_ASR_BACKEND.

When an engine reports itself unavailable, expand its row's Why? panel and use Learn more to jump straight to that engine's page here. The row's own message stays deliberately generic — an availability probe can carry local paths or credentials, so it is never shown verbatim — and the page below is where the actual requirements and setup steps live.

The compute device (CUDA/ROCm/MPS/CPU) is auto-detected; pin it under Settings → Performance & Device (or OMNIVOICE_DEVICE) if auto-detect picks wrong — see performance.

Measured speed/VRAM numbers live in benchmarks; what each engine can do expressively in expressive-speech; sidecar disk footprints in disk-usage; the bar a new engine must clear in engine-acceptance.

New to VoiceStudio? Install the app first — macOS (first launch needs the one-time right-click → Open Gatekeeper approval), Windows, Linux, Docker.

Text-to-speech

Engine Guide Runs on Cloning Enabled by
VoiceStudio (OmniVoice) — default omnivoice CUDA · MPS · CPU ✅ installed by default
VoxCPM2 voxcpm2 CUDA · MPS · CPU ✅ + voice design pip install "voxcpm>=2.0.3"
MOSS-TTS-Nano moss-tts-nano CUDA · CPU ✅ (ref only) clone + uv pip install -e .
KittenTTS kittentts CPU — (8 preset voices) pip install kittentts
MLX-Audio (Kokoro, CSM, Dia, …) mlx-audio Apple Silicon model-dependent pip install mlx-audio
CosyVoice 3 cosyvoice CUDA · CPU ✅ clone + requirements
GPT-SoVITS gpt-sovits external server ✅ its own API server
Sherpa-ONNX sherpa-onnx CUDA · CPU — pip install sherpa-onnx + model dir
IndexTTS 2.5 indextts CUDA · CPU ✅ + emotion one-click sidecar install
OmniVoice GGUF omnivoice-gguf CUDA · MPS · CPU ✅ bundled binary
Supertonic-3 supertonic3 CPU — (7 preset voices) uv sync --extra supertonic + license
MOSS-TTS-v1.5 (8B) moss-tts-v15 CUDA · CPU ✅ clone + env var
dots.tts (2B) dots-tts CUDA · CPU (not Windows) ✅ clone + env var
OmniVoice (subprocess) omnivoice-subprocess CUDA · MPS · CPU ✅ opt-in pick off MPS; automatic via default OmniVoice on MPS
PocketTTS (Kyutai) pockettts CPU (not Intel Mac) ✅ uv sync --extra pockettts + license
Confucius4-TTS confucius4-tts CUDA · CPU ✅ clone + env var
audio.cpp (Breeze-TTS-2) audio-cpp CPU + Vulkan/Metal/CUDA/HIP/ROCm where compiled ✅ + voice design prebuilt binary + env var (weights research/non-commercial)

Speech-to-text

Engine Guide Runs on Best at Enabled by
WhisperX whisperx CUDA · CPU dubbing (word timestamps + diarization) installed by default
Faster-Whisper faster-whisper CUDA · CPU general transcription installed by default
Faster-Whisper (isolated) faster-whisper-isolated CUDA · CPU unattended batches opt-in pick
MLX Whisper mlx-whisper Apple Silicon Mac default pip install mlx-whisper
PyTorch Whisper pytorch-whisper CUDA · MPS · CPU ROCm hosts installed by default
Parakeet TDT (NeMo) nemo-parakeet CUDA · CPU 25 languages, fast CPU separate venv (never the app's)
Parakeet TDT (MLX) parakeet-mlx Apple Silicon dictation, 25 EU languages default on mac-ARM source installs
Moonshine moonshine CPU edge/low-power, no timestamps pip install (see guide)
FunASR (SenseVoice) funasr CUDA · CPU 50+ languages, inline diarization pip install funasr
Sherpa-ONNX dictation sherpa-onnx-asr CPU live streaming dictation curated model download
OpenAI-compatible (local or remote) openai-compatible-asr network a configured endpoint; loopback stays local Model Catalogue

Speaker diarization is not an engine registry of its own — the dub pipeline uses pyannote (HF-gated; see diarization) and FunASR can diarize inline with its cam++ speaker model.