* docs(engines): a guide for every engine + index; fix two engine-metadata bugs 21 new pages under docs/engines/ (10 TTS, 10 ASR, index README) — every registered engine now has one: what it's for, platform support, model env vars, quirks with issue refs. Linked from both READMEs' engine sections. Code fixes found while verifying facts against the registries: - KittenTTS docstring claimed default voice 'Jasper'; the code default is expr-voice-2-f - the isolated-ASR sidecar read only ASR_MODEL_FW while the download preflight read ASR_MODEL_FASTER — set one and the other quietly used a different model; both now resolve ASR_MODEL_FW-override → ASR_MODEL_FASTER - moonshine's install hint named 'useful-moonshine', a package the backend never imports; now moonshine-onnx / moonshine-voice Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(changelog): entries for the engine guides + sidecar model fix (#1556) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(engines): second-harvest fixes — TLS guidance, matrix/code alignment, CN counts - README matrix aligned to gpu_compat (the code is the source of truth): CosyVoice macOS is CPU not MPS, IndexTTS and GGUF gain their real CUDA/CPU/MPS cells - gpt-sovits guide: prefer https/tunnel for non-loopback servers, plaintext warning; first-use download guidance on both OmniVoice pages - preflight empty-env fallback matches the sidecar (ASR_MODEL_FASTER='' no longer resolves a different repo) - nano installs via uv pip; kitten log level wording; index links install guides incl. the Gatekeeper step; README_CN engine counts 16/11 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(readme-cn): the all-engines-local claim now excludes the remote client Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Engine guides
One page per engine: what it's for, what it needs, how to enable it, and its
quirks. Select engines in Model Catalogue → Engines (or quick-switch with
Ctrl/Cmd+E), or pin one with
OMNIVOICE_TTS_BACKEND / OMNIVOICE_ASR_BACKEND.
Measured speed/VRAM numbers live in benchmarks; what each engine can do expressively in expressive-speech; sidecar disk footprints in disk-usage; the bar a new engine must clear in engine-acceptance.
New to VoiceStudio? Install the app first — macOS (first launch needs the one-time right-click → Open Gatekeeper approval), Windows, Linux, Docker.
Text-to-speech
| Engine | Guide | Runs on | Cloning | Enabled by |
|---|---|---|---|---|
| VoiceStudio (OmniVoice) — default | omnivoice | CUDA · MPS · CPU | ✅ | installed by default |
| VoxCPM2 | voxcpm2 | CUDA · MPS · CPU | ✅ + voice design | pip install "voxcpm>=2.0.3" |
| MOSS-TTS-Nano | moss-tts-nano | CUDA · CPU | ✅ (ref only) | clone + uv pip install -e . |
| KittenTTS | kittentts | CPU | — (8 preset voices) | pip install kittentts |
| MLX-Audio (Kokoro, CSM, Dia, …) | mlx-audio | Apple Silicon | model-dependent | pip install mlx-audio |
| CosyVoice 3 | cosyvoice | CUDA · CPU | ✅ | clone + requirements |
| GPT-SoVITS | gpt-sovits | external server | ✅ | its own API server |
| Sherpa-ONNX | sherpa-onnx | CUDA · CPU | — | pip install sherpa-onnx + model dir |
| IndexTTS 2.5 | indextts | CUDA · CPU | ✅ + emotion | one-click sidecar install |
| OmniVoice GGUF | omnivoice-gguf | CUDA · MPS · CPU | ✅ | bundled binary |
| Supertonic-3 | supertonic3 | CPU | — (7 preset voices) | uv sync --extra supertonic + license |
| MOSS-TTS-v1.5 (8B) | moss-tts-v15 | CUDA · CPU | ✅ | clone + env var |
| dots.tts (2B) | dots-tts | CUDA · CPU (not Windows) | ✅ | clone + env var |
| OmniVoice (subprocess) | omnivoice-subprocess | CUDA · MPS · CPU | ✅ | opt-in pick, no install |
| PocketTTS (Kyutai) | pockettts | CPU (not Intel Mac) | ✅ | uv sync --extra pockettts + license |
| Confucius4-TTS | confucius4-tts | CUDA · CPU | ✅ | clone + env var |
Speech-to-text
| Engine | Guide | Runs on | Best at | Enabled by |
|---|---|---|---|---|
| WhisperX | whisperx | CUDA · CPU | dubbing (word timestamps + diarization) | installed by default |
| Faster-Whisper | faster-whisper | CUDA · CPU | general transcription | installed by default |
| Faster-Whisper (isolated) | faster-whisper-isolated | CUDA · CPU | unattended batches | opt-in pick |
| MLX Whisper | mlx-whisper | Apple Silicon | Mac default | pip install mlx-whisper |
| PyTorch Whisper | pytorch-whisper | CUDA · MPS · CPU | ROCm hosts | installed by default |
| Parakeet TDT (NeMo) | nemo-parakeet | CUDA · CPU | 25 languages, fast CPU | separate venv (never the app's) |
| Parakeet TDT (MLX) | parakeet-mlx | Apple Silicon | dictation, 25 EU languages | default on mac-ARM source installs |
| Moonshine | moonshine | CPU | edge/low-power, no timestamps | pip install (see guide) |
| FunASR (SenseVoice) | funasr | CUDA · CPU | 50+ languages, inline diarization | pip install funasr |
| Sherpa-ONNX dictation | sherpa-onnx-asr | CPU | live streaming dictation | curated model download |
| OpenAI-compatible (remote) | openai-compatible-asr | network | offloading to a server (audio leaves the machine) | Model Catalogue |
Speaker diarization is not an engine registry of its own — the dub pipeline
uses pyannote (HF-gated; see diarization) and
FunASR can diarize inline with its cam++ speaker model.