+
+
+ 🎙️ Voice Cloning
+ 3-second clip → mirror any voice. 646 languages, zero-shot.
+ |
+
+ 🎨 Voice Design
+ Gender, age, accent, pitch, speed, emotion, dialect — dial it in.
+ |
+
+ 🎬 Video Dubbing
+ YouTube URL or file → transcribe → translate → re-voice → MP4.
+ |
+
+ 📖 Audiobook Editor
+ Import text, EPUB, or PDF. Auto-chapter, loudnorm, metadata. Export .m4b.
+ |
+
+
+
+ 🎭 Stories
+ Multi-voice editor. Assign voices per-line, preview, export full cast.
+ |
+
+ ⌨️ Dictation Widget
+ ⌘+⇧+Space from any app. Transcribes, auto-pastes, disappears.
+ |
+
+ 🔐 100% Local
+ No keys, no cloud, no accounts. Your machine only.
+ |
+
+ 🤖 MCP Server
+ Use OmniVoice from Claude, Cursor, or any MCP client.
+ |
+
+
+
+
@@ -296,7 +284,12 @@ OmniVoice Studio gives you professional-grade AI tools without the subscription
### 🗣️ TTS Engines
-OmniVoice ships a multi-engine TTS backend. The default engine (OmniVoice) is always available; additional engines are opt-in and auto-detected. Switch engines in **Settings → TTS Engine** or via the `OMNIVOICE_TTS_BACKEND` env var.
+**14 engines, one picker.** OmniVoice (default, 600+ languages) is always available; CosyVoice 3, GPT-SoVITS, VoxCPM2, MOSS-TTS-Nano, KittenTTS, MLX-Audio, and Sherpa-ONNX are opt-in and auto-detected — plus six lazy-installed heavyweights (IndexTTS 2, OmniVoice GGUF, Supertonic 3, MOSS-TTS-v1.5, dots.tts, Confucius4-TTS). Switch in **Settings → TTS Engine** or via the `OMNIVOICE_TTS_BACKEND` env var.
+
+
+📊 The full matrix — 14 engines × platform × clone/instruct × license
+
+
| Engine | Languages | Clone | Instruct | Linux | macOS ARM | Windows | License |
|--------|:---------:|:-----:|:--------:|:-----:|:---------:|:-------:|:-------:|
@@ -319,11 +312,18 @@ OmniVoice ships a multi-engine TTS backend. The default engine (OmniVoice) is al
>
> **MOSS-TTS-v1.5** (8B, ~16 GB weights) and **dots.tts** (2B, ~9 GB weights) are heavyweight opt-in engines that run in their own isolated venv from a local clone — see [MOSS-TTS-v1.5](docs/engines/moss-tts-v15.md) and [dots.tts](docs/engines/dots-tts.md). Neither claims Apple-Silicon **MPS** (upstream is CUDA/CPU only; on a Mac they run on CPU). dots.tts upstream is Linux/macOS only — no Windows path. **Confucius4-TTS** (14-language cross-lingual zero-shot cloning) is similar — its own Python 3.10 venv from a clone; CUDA recommended, CPU validated end-to-end (slow, ~17× realtime; no MPS — tested slower than CPU); see [Confucius4-TTS](docs/engines/confucius4-tts.md).
+
+
### 🎧 ASR Engines
-OmniVoice ships a multi-engine ASR (speech-to-text) backend that powers dictation, video dubbing, and subtitle generation — all fully local. **WhisperX** is the cross-platform default; the rest are opt-in and auto-detected. Switch in **Settings → ASR Engine** or via the `OMNIVOICE_ASR_BACKEND` env var.
+**9 engines, all fully local** — they power dictation, video dubbing, and subtitles. **WhisperX** is the cross-platform default (~100 languages, word-level timing); the rest are opt-in and auto-detected. Switch in **Settings → ASR Engine** or via the `OMNIVOICE_ASR_BACKEND` env var.
+
+
+📊 The full lineup — 9 engines, what each is best at, and compute-type notes
+
+
| Engine | `OMNIVOICE_ASR_BACKEND` | Languages | Best for |
|--------|-------------------------|:---------:|----------|
@@ -341,6 +341,8 @@ OmniVoice ships a multi-engine ASR (speech-to-text) backend that powers dictatio
> **GPU without efficient float16?** On older NVIDIA GPUs (Maxwell/Pascal, GTX 16xx) or after a CTranslate2/cuDNN mismatch, the CTranslate2 ASR engines (WhisperX, Faster-Whisper) can't run `float16` and OmniVoice automatically retries on `int8` — no config needed. If transcription still fails, pin the compute type with the `ASR_COMPUTE_TYPE` env var (escape hatch): `ASR_COMPUTE_TYPE=int8` (or `float32` for CPU). Set it to `int8` and restart the backend.
+
+
---
## 🏗️ Architecture
@@ -361,11 +363,50 @@ OmniVoice ships a multi-engine ASR (speech-to-text) backend that powers dictatio
CUDA / MPS / ROCm / CPU (auto-detected + routed)
```
+
+
+## 🔌 OpenAI-compatible API
+
+Already have a script, agent, or tool that speaks OpenAI's audio API? Point it at `http://localhost:3900/v1` — no key needed, no code changes. The backend ships a drop-in surface for the audio endpoints, wired to whichever TTS/ASR engine you have active (and yes, `voice` accepts your cloned voice-profile IDs).
+
+| Endpoint | What it does |
+|---|---|
+| `POST /v1/audio/speech` | TTS — text in; `mp3` / `wav` / `flac` / `opus` / `pcm` out. `tts-1` / `tts-1-hd` map to your active engine; OpenAI voice names (`alloy`, …) are accepted. |
+| `POST /v1/audio/transcriptions` | STT — audio file in; `json`, `text`, `verbose_json`, `srt`, or `vtt` out. `whisper-1` maps to your active ASR engine. |
+| `GET /v1/audio/voices` | OmniVoice extension — lists every voice profile and engine, so clients can discover your clones. |
+
+```sh
+curl http://localhost:3900/v1/audio/speech \
+ -H "Content-Type: application/json" \
+ -d '{"model": "tts-1", "voice": "alloy", "input": "Generated on my own hardware.", "response_format": "wav"}' \
+ --output speech.wav
+```
+
+```python
+from openai import OpenAI
+client = OpenAI(base_url="http://localhost:3900/v1", api_key="none") # any string works — nothing checks it
+
+result = client.audio.transcriptions.create(model="whisper-1", file=open("clip.wav", "rb"))
+print(result.text)
+```
+
+Want the whole surface (100+ endpoints)? The full REST API reference is embedded in the app — **Settings → OpenAPI Reference** (Scalar-powered), or the `{}` button in the footer.
+
---
## 🗺️ Roadmap
-### ✅ Shipped
+### 🔜 Up Next
+
+- 🎬 **Lip-sync v2** — visual speech timing with wav2lip
+- 🌐 **Hosted Demo** — try OmniVoice without installing anything
+- 🔌 **Plugin Marketplace** — community-contributed TTS engines and effects
+- 🎵 **Real-time Voice Changer** — live microphone transformation during calls
+
+
+✅ Everything shipped so far — the receipts, by category
+
+
| Category | Features |
|----------|----------|
@@ -389,12 +430,7 @@ OmniVoice ships a multi-engine ASR (speech-to-text) backend that powers dictatio
| **Remote Backend** | Point the desktop UI at a remote backend URL with bearer auth (Tailscale-documented) |
| **Reliability** | Stall watchdog on bootstrap splash, per-engine GPU compatibility matrix, actionable errors for non-executable engine binaries, setuptools auto-repair |
-### 🔜 Up Next
-
-- 🎬 **Lip-sync v2** — visual speech timing with wav2lip
-- 🌐 **Hosted Demo** — try OmniVoice without installing anything
-- 🔌 **Plugin Marketplace** — community-contributed TTS engines and effects
-- 🎵 **Real-time Voice Changer** — live microphone transformation during calls
+
---
@@ -445,8 +481,13 @@ OmniVoice is **free** and **AGPL-3.0** — no paid tier, no SaaS revenue. Sponso

+
+
We respond to setup questions within hours, not days.
+
+What happens in there
+
| Channel | What happens there |
@@ -457,7 +498,7 @@ OmniVoice is **free** and **AGPL-3.0** — no paid tier, no SaaS revenue. Sponso
| `#dev` | Architecture discussions, PR reviews, engine integrations |
| `#announcements` | Release notes, breaking changes, early access |
-**[→ Join the Discord](https://discord.gg/bzQavDfVV9)** — we respond to setup questions within hours, not days.
+
---
@@ -465,7 +506,7 @@ OmniVoice is **free** and **AGPL-3.0** — no paid tier, no SaaS revenue. Sponso
## 🤝 Contributing
-We welcome contributions of all kinds — bug fixes, new TTS engine adapters, UI improvements, docs, and translations.
+Yes please — bug fixes, new TTS engine adapters, UI improvements, docs, translations. All of it.
- 📖 Read the **[Contributing Guide](CONTRIBUTING.md)** for setup, code style, and PR workflow
- 🐛 Browse [good first issues](https://github.com/debpalash/OmniVoice-Studio/labels/good%20first%20issue)
diff --git a/docs/screenshot-dub.png b/docs/screenshot-dub.png
index 2c7ce175..041a548a 100644
Binary files a/docs/screenshot-dub.png and b/docs/screenshot-dub.png differ