diff --git a/README.md b/README.md index 872e4c9d..75242f9c 100644 --- a/README.md +++ b/README.md @@ -10,6 +10,7 @@ Why OVS · TTS Engines · ASR Engines · + API · Sponsors · Donate · Contributing · @@ -34,8 +35,18 @@ OmniVoice Studio — Launchpad +> **Your voice is the most personal data you have. So why rent it back from a cloud?** Every mainstream voice tool ships your audio to someone else's server and bills you monthly for the privilege. OmniVoice Studio flips that: clone, design, dub, and dictate on your own hardware — 646 languages, no meter running, nothing leaving your machine. + +
+ +| 🔑 No API keys | 🙅 No accounts | ☁️ No cloud | 💳 No subscription | +|:---:|:---:|:---:|:---:| +| nothing to paste in | nothing to sign up for | your audio stays home | it's your computer | + +
+ > [!WARNING] -> **OmniVoice Studio is in active beta.** Things may break between releases. For the latest features and fixes, clone the repo and run from source rather than using pre-built installers. Bug reports and PRs are very welcome — [open an issue](https://github.com/debpalash/OmniVoice-Studio/issues) or [join Discord](https://discord.gg/bzQavDfVV9). +> **OmniVoice Studio is in active beta.** Things may break between releases — for the latest features and fixes, clone the repo and run from source rather than the pre-built installers. Bug reports and PRs are very welcome: [open an issue](https://github.com/debpalash/OmniVoice-Studio/issues) or [join Discord](https://discord.gg/bzQavDfVV9).

@@ -47,142 +58,6 @@
- - -## ✨ Features - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -
-

🎙️ Voice Cloning

-

3-second clip → mirror any voice.
646 languages, zero-shot.

-
-

🎨 Voice Design

-

Gender, age, accent, pitch, speed,
emotion, dialect — dial it in.

-
-

🎬 Video Dubbing

-

YouTube URL or file → transcribe →
translate → re-voice → MP4.

-
-

📖 Audiobook Editor

-

Import text, EPUB, or PDF. Auto-chapter,
loudnorm, metadata. Export .m4b.

-
-

🎭 Stories

-

Multi-voice editor. Assign voices
per-line, preview, export full cast.

-
-

⌨️ Dictation Widget

-

⌘+⇧+Space from any app.
Transcribes, auto-pastes, disappears.

-
-

🔊 Vocal Isolation

-

Demucs-powered. Splits speech
from music, keeps the background.

-
-

👥 Speaker Diarization

-

Pyannote + WhisperX.
Auto-identifies who said what.

-
-

📦 Batch Queue

-

Drop 50 videos, walk away.
Progress bars per job.

-
-

🤖 MCP Server

-

Use OmniVoice from Claude,
Cursor, or any MCP client.

-
-

🛡️ AI Watermark

-

AudioSeal (Meta). Invisible,
survives compression.

-
-

🔬 Diagnostics

-

Self-check, error journal,
scrubbed diagnostic bundle.

-
-

🔐 100% Local

-

No keys, no cloud, no accounts.
Your machine only.

-
-

⚡ GPU Auto-Detect

-

CUDA · MPS · ROCm · CPU.
≤8 GB? Auto-offloads.

-
-

🧩 Extensible

-

Subclass TTSbackend,
add any engine in ~50 lines.

-
-

🧭 Engine Routing

-

Preflight GPU check per engine.
No silent CPU fallback.

-
-

🎒 Portable Personas

-

Export voices as .ovsvoice
bundles — identity + watermark.

-
-

♾️ Unlimited TTS

-

Sentence-chunked generation.
No length cap. Streaming via WS.

-
-

🌐 Remote Backend

-

Point UI at a remote server.
Tailscale-friendly. Bearer auth.

-
-

🧠 Dictation + LLM

-

Local LLM cleanup of transcripts.
Optional echo cancellation.

-
- ---- - - - -## ⚡ Quickstart - -
- Download macOS DMG - Download Windows MSI - Download Linux AppImage - Download Debian .deb -
- macOS: first launch needs a one-time approval — right-click → Open (or System Settings → Privacy & Security → "Open Anyway" on macOS 15). No Terminal needed. Why? -
- Intel Macs are not supported for the local backend: the app UI installs, but the Python backend cannot run because PyTorch no longer ships Intel-Mac (x86_64) wheels (#889) — see docs/install/macos.md. -
- -Per-OS install guides — pick yours and follow it end-to-end: - -- **macOS** — [docs/install/macos.md](docs/install/macos.md) -- **Windows** — [docs/install/windows.md](docs/install/windows.md) -- **Linux** — [docs/install/linux.md](docs/install/linux.md) -- **Docker** — [docs/install/docker.md](docs/install/docker.md) · [Docker Hub: `palashdeb/omnivoice-studio`](https://hub.docker.com/r/palashdeb/omnivoice-studio) - -Stuck? Run the built-in self-check first — **Settings → About → "Run -self-check"** in the app, or `uv run python backend/main.py --diagnose` from -a checkout (`--deep` also test-loads the active engine). Then see -[docs/install/troubleshooting.md](docs/install/troubleshooting.md) for the -top 10 install errors. The in-app error UI deeplinks to those entries when -something breaks at runtime, and **Settings → About → "Save diagnostic -bundle"** packages scrubbed logs + the self-check report for bug reports. - -For Hugging Face token setup, see -[docs/setup/huggingface-token.md](docs/setup/huggingface-token.md). For -diarization-specific gating, see -[docs/features/diarization.md](docs/features/diarization.md). For download -speed, the ⚡ fast-download (Xet) status, and restricted-network / mirror -options, see [docs/downloading-models.md](docs/downloading-models.md). - ## 📸 See it in action @@ -209,7 +84,7 @@ options, see [docs/downloading-models.md](docs/downloading-models.md). Video Dubbing
Video Dubbing
- Upload a file or paste a YouTube URL → transcribe, translate, re-voice, export MP4. + A real dub, end to end: 37 segments transcribed, translated to Bengali, re-voiced, and timed — ready to export as MP4. @@ -240,6 +115,119 @@ options, see [docs/downloading-models.md](docs/downloading-models.md). --- + + +## ✨ Features + +The eight headliners — and twelve more waiting under the fold. + + + + + + + + + + + + + + +
+

🎙️ Voice Cloning

+

3-second clip → mirror any voice.
646 languages, zero-shot.

+
+

🎨 Voice Design

+

Gender, age, accent, pitch, speed,
emotion, dialect — dial it in.

+
+

🎬 Video Dubbing

+

YouTube URL or file → transcribe →
translate → re-voice → MP4.

+
+

📖 Audiobook Editor

+

Import text, EPUB, or PDF. Auto-chapter,
loudnorm, metadata. Export .m4b.

+
+

🎭 Stories

+

Multi-voice editor. Assign voices
per-line, preview, export full cast.

+
+

⌨️ Dictation Widget

+

⌘+⇧+Space from any app.
Transcribes, auto-pastes, disappears.

+
+

🔐 100% Local

+

No keys, no cloud, no accounts.
Your machine only.

+
+

🤖 MCP Server

+

Use OmniVoice from Claude,
Cursor, or any MCP client.

+
+ +
+…and 12 more — isolation, diarization, batch, watermarking, diagnostics, and friends + +
+ +- 🔊 **Vocal Isolation** — Demucs-powered: splits speech from music and keeps the background bed. +- 👥 **Speaker Diarization** — Pyannote + WhisperX auto-identify who said what. +- 📦 **Batch Queue** — drop 50 videos, walk away; per-job progress bars. +- 🛡️ **AI Watermark** — AudioSeal (Meta): invisible, survives compression. +- 🔬 **Diagnostics** — self-check suite, error journal, scrubbed diagnostic bundles. +- ⚡ **GPU Auto-Detect** — CUDA · MPS · ROCm · CPU; ≤8 GB VRAM auto-offloads. +- 🧭 **Engine routing** — preflight GPU check per engine; no silent CPU fallback. +- 🧩 **Extensible** — subclass `TTSBackend`, add any engine in ~50 lines. +- 🎒 **Portable personas** — export voices as `.ovsvoice` bundles: identity + watermark. +- ♾️ **Unlimited TTS** — sentence-chunked generation, no length cap, streaming via WebSocket. +- 🌐 **Remote backend** — point the UI at a remote server; Tailscale-friendly, bearer auth. +- 🧠 **Dictation + LLM** — local-LLM cleanup of transcripts, optional echo cancellation. + +
+ +--- + + + +## ⚡ Quickstart + +
+ Download macOS DMG + Download Windows MSI + Download Linux AppImage + Download Debian .deb +
+ macOS: first launch needs a one-time approval — right-click → Open (or System Settings → Privacy & Security → "Open Anyway" on macOS 15). No Terminal needed. Why? +
+ Intel Macs are not supported for the local backend: the app UI installs, but the Python backend cannot run because PyTorch no longer ships Intel-Mac (x86_64) wheels (#889) — see docs/install/macos.md. +
+ +Pick your OS and follow the guide end-to-end: + +- 🍎 **macOS** — [docs/install/macos.md](docs/install/macos.md) +- 🪟 **Windows** — [docs/install/windows.md](docs/install/windows.md) +- 🐧 **Linux** — [docs/install/linux.md](docs/install/linux.md) +- 🐳 **Docker** — [docs/install/docker.md](docs/install/docker.md) · [Docker Hub: `palashdeb/omnivoice-studio`](https://hub.docker.com/r/palashdeb/omnivoice-studio) + +
+🧰 Stuck? Self-checks, tokens & restricted networks + +
+ +Run the built-in self-check first — **Settings → About → "Run +self-check"** in the app, or `uv run python backend/main.py --diagnose` from +a checkout (`--deep` also test-loads the active engine). Then see +[docs/install/troubleshooting.md](docs/install/troubleshooting.md) for the +top 10 install errors. The in-app error UI deeplinks to those entries when +something breaks at runtime, and **Settings → About → "Save diagnostic +bundle"** packages scrubbed logs + the self-check report for bug reports. + +For Hugging Face token setup, see +[docs/setup/huggingface-token.md](docs/setup/huggingface-token.md). For +diarization-specific gating, see +[docs/features/diarization.md](docs/features/diarization.md). For download +speed, the ⚡ fast-download (Xet) status, and restricted-network / mirror +options, see [docs/downloading-models.md](docs/downloading-models.md). + +
+ +--- + ## 💡 Why OmniVoice? @@ -264,7 +252,7 @@ ElevenLabs charges **$5–$330/mo** and processes your audio on their servers. O | **Self-check** | ❌ | ✅ Diagnostics suite, error journal, scrubbed debug bundles | | **Customizable** | ❌ Closed | ✅ Fork it, extend it, ship it | -OmniVoice Studio gives you professional-grade AI tools without the subscription or the cloud. +Professional-grade voice AI, minus the subscription and the cloud.

@@ -296,7 +284,12 @@ OmniVoice Studio gives you professional-grade AI tools without the subscription ### 🗣️ TTS Engines -OmniVoice ships a multi-engine TTS backend. The default engine (OmniVoice) is always available; additional engines are opt-in and auto-detected. Switch engines in **Settings → TTS Engine** or via the `OMNIVOICE_TTS_BACKEND` env var. +**14 engines, one picker.** OmniVoice (default, 600+ languages) is always available; CosyVoice 3, GPT-SoVITS, VoxCPM2, MOSS-TTS-Nano, KittenTTS, MLX-Audio, and Sherpa-ONNX are opt-in and auto-detected — plus six lazy-installed heavyweights (IndexTTS 2, OmniVoice GGUF, Supertonic 3, MOSS-TTS-v1.5, dots.tts, Confucius4-TTS). Switch in **Settings → TTS Engine** or via the `OMNIVOICE_TTS_BACKEND` env var. + +
+📊 The full matrix — 14 engines × platform × clone/instruct × license + +
| Engine | Languages | Clone | Instruct | Linux | macOS ARM | Windows | License | |--------|:---------:|:-----:|:--------:|:-----:|:---------:|:-------:|:-------:| @@ -319,11 +312,18 @@ OmniVoice ships a multi-engine TTS backend. The default engine (OmniVoice) is al > > **MOSS-TTS-v1.5** (8B, ~16 GB weights) and **dots.tts** (2B, ~9 GB weights) are heavyweight opt-in engines that run in their own isolated venv from a local clone — see [MOSS-TTS-v1.5](docs/engines/moss-tts-v15.md) and [dots.tts](docs/engines/dots-tts.md). Neither claims Apple-Silicon **MPS** (upstream is CUDA/CPU only; on a Mac they run on CPU). dots.tts upstream is Linux/macOS only — no Windows path. **Confucius4-TTS** (14-language cross-lingual zero-shot cloning) is similar — its own Python 3.10 venv from a clone; CUDA recommended, CPU validated end-to-end (slow, ~17× realtime; no MPS — tested slower than CPU); see [Confucius4-TTS](docs/engines/confucius4-tts.md). +
+ ### 🎧 ASR Engines -OmniVoice ships a multi-engine ASR (speech-to-text) backend that powers dictation, video dubbing, and subtitle generation — all fully local. **WhisperX** is the cross-platform default; the rest are opt-in and auto-detected. Switch in **Settings → ASR Engine** or via the `OMNIVOICE_ASR_BACKEND` env var. +**9 engines, all fully local** — they power dictation, video dubbing, and subtitles. **WhisperX** is the cross-platform default (~100 languages, word-level timing); the rest are opt-in and auto-detected. Switch in **Settings → ASR Engine** or via the `OMNIVOICE_ASR_BACKEND` env var. + +
+📊 The full lineup — 9 engines, what each is best at, and compute-type notes + +
| Engine | `OMNIVOICE_ASR_BACKEND` | Languages | Best for | |--------|-------------------------|:---------:|----------| @@ -341,6 +341,8 @@ OmniVoice ships a multi-engine ASR (speech-to-text) backend that powers dictatio > **GPU without efficient float16?** On older NVIDIA GPUs (Maxwell/Pascal, GTX 16xx) or after a CTranslate2/cuDNN mismatch, the CTranslate2 ASR engines (WhisperX, Faster-Whisper) can't run `float16` and OmniVoice automatically retries on `int8` — no config needed. If transcription still fails, pin the compute type with the `ASR_COMPUTE_TYPE` env var (escape hatch): `ASR_COMPUTE_TYPE=int8` (or `float32` for CPU). Set it to `int8` and restart the backend. +
+ --- ## 🏗️ Architecture @@ -361,11 +363,50 @@ OmniVoice ships a multi-engine ASR (speech-to-text) backend that powers dictatio CUDA / MPS / ROCm / CPU (auto-detected + routed) ``` + + +## 🔌 OpenAI-compatible API + +Already have a script, agent, or tool that speaks OpenAI's audio API? Point it at `http://localhost:3900/v1` — no key needed, no code changes. The backend ships a drop-in surface for the audio endpoints, wired to whichever TTS/ASR engine you have active (and yes, `voice` accepts your cloned voice-profile IDs). + +| Endpoint | What it does | +|---|---| +| `POST /v1/audio/speech` | TTS — text in; `mp3` / `wav` / `flac` / `opus` / `pcm` out. `tts-1` / `tts-1-hd` map to your active engine; OpenAI voice names (`alloy`, …) are accepted. | +| `POST /v1/audio/transcriptions` | STT — audio file in; `json`, `text`, `verbose_json`, `srt`, or `vtt` out. `whisper-1` maps to your active ASR engine. | +| `GET /v1/audio/voices` | OmniVoice extension — lists every voice profile and engine, so clients can discover your clones. | + +```sh +curl http://localhost:3900/v1/audio/speech \ + -H "Content-Type: application/json" \ + -d '{"model": "tts-1", "voice": "alloy", "input": "Generated on my own hardware.", "response_format": "wav"}' \ + --output speech.wav +``` + +```python +from openai import OpenAI +client = OpenAI(base_url="http://localhost:3900/v1", api_key="none") # any string works — nothing checks it + +result = client.audio.transcriptions.create(model="whisper-1", file=open("clip.wav", "rb")) +print(result.text) +``` + +Want the whole surface (100+ endpoints)? The full REST API reference is embedded in the app — **Settings → OpenAPI Reference** (Scalar-powered), or the `{}` button in the footer. + --- ## 🗺️ Roadmap -### ✅ Shipped +### 🔜 Up Next + +- 🎬 **Lip-sync v2** — visual speech timing with wav2lip +- 🌐 **Hosted Demo** — try OmniVoice without installing anything +- 🔌 **Plugin Marketplace** — community-contributed TTS engines and effects +- 🎵 **Real-time Voice Changer** — live microphone transformation during calls + +
+✅ Everything shipped so far — the receipts, by category + +
| Category | Features | |----------|----------| @@ -389,12 +430,7 @@ OmniVoice ships a multi-engine ASR (speech-to-text) backend that powers dictatio | **Remote Backend** | Point the desktop UI at a remote backend URL with bearer auth (Tailscale-documented) | | **Reliability** | Stall watchdog on bootstrap splash, per-engine GPU compatibility matrix, actionable errors for non-executable engine binaries, setuptools auto-repair | -### 🔜 Up Next - -- 🎬 **Lip-sync v2** — visual speech timing with wav2lip -- 🌐 **Hosted Demo** — try OmniVoice without installing anything -- 🔌 **Plugin Marketplace** — community-contributed TTS engines and effects -- 🎵 **Real-time Voice Changer** — live microphone transformation during calls +
--- @@ -445,8 +481,13 @@ OmniVoice is **free** and **AGPL-3.0** — no paid tier, no SaaS revenue. Sponso
Join Discord +
+ We respond to setup questions within hours, not days.
+
+What happens in there +
| Channel | What happens there | @@ -457,7 +498,7 @@ OmniVoice is **free** and **AGPL-3.0** — no paid tier, no SaaS revenue. Sponso | `#dev` | Architecture discussions, PR reviews, engine integrations | | `#announcements` | Release notes, breaking changes, early access | -**[→ Join the Discord](https://discord.gg/bzQavDfVV9)** — we respond to setup questions within hours, not days. +
--- @@ -465,7 +506,7 @@ OmniVoice is **free** and **AGPL-3.0** — no paid tier, no SaaS revenue. Sponso ## 🤝 Contributing -We welcome contributions of all kinds — bug fixes, new TTS engine adapters, UI improvements, docs, and translations. +Yes please — bug fixes, new TTS engine adapters, UI improvements, docs, translations. All of it. - 📖 Read the **[Contributing Guide](CONTRIBUTING.md)** for setup, code style, and PR workflow - 🐛 Browse [good first issues](https://github.com/debpalash/OmniVoice-Studio/labels/good%20first%20issue) diff --git a/docs/screenshot-dub.png b/docs/screenshot-dub.png index 2c7ce175..041a548a 100644 Binary files a/docs/screenshot-dub.png and b/docs/screenshot-dub.png differ