OmniVoice Logo

OmniVoice Studio

Your Local Cinematic AI Dubbing Studio

FeaturesGetting StartedRoadmapChangelog


OmniVoice Studio Design Interface
High-density Voice Design & Cloning workspace.

Local, full-stack voice generation and cinematic dubbing. No API keys. No cloud. Just run it. Built on the open-source OmniVoice 600-language zero-shot diffusion model.

Features

  • 🎬 Video Dubbing — transcribe, translate, re-voice, and mux back into MP4 dynamically.
  • 🎧 Vocal Isolation — built-in demucs automatically splits speech from music, keeping original background audio perfectly preserved.
  • 🧬 Voice Cloning & Design — Clone specific voices from just a 3-second audio clip, or design completely new studio profiles with tags like female, british accent, excited.
  • Cross-Platform Native Execution — Auto-detects and accelerates inference using Apple Silicon (MPS), NVIDIA (CUDA), AMD (ROCm), or standard CPU.

OmniVoice Studio Dubbing Interface
The timeline-based cinematic dubbing studio.

🚀 Getting Started

Quickly get OmniVoice Studio running locally on your hardware.

Prerequisites: Ensure ffmpeg is installed on your system. Install standard web tooling: Bun and uv.

git clone https://github.com/debpalash/OmniVoice-Studio.git
cd OmniVoice-Studio

# Backend
uv sync

# Frontend
bun install
bun dev

OmniVoice Studio launches exactly two micro-services:

Service Protocol Details
Frontend http://localhost:5173 The real-time React UI — spanning cloning, design, and audio workspace.
Backend http://localhost:8000 The FastAPI server handling model inference, translation pipelines, transcriber tasks.

Note

First run optimization: Model weights (approx. 1.2 GB) automatically download from HuggingFace the first time you execute a generation sequence. Subsequent launches trigger instantly from cache. (Tip: Set HF_TOKEN in your environment for faster, authenticated downloads!)


🗺️ Roadmap

The studio is highly functional today, but we are aggressively expanding. Watch the roadmap to see what's shipping next:

🌟 Completed Milestones

  • Zero-shot voice cloning & complex voice design.
  • Full video cinematic dubbing pipeline (transcribe → translate → synthesize → mux).
  • Vocal isolation utilizing demucs alongside background audio retention.
  • Embedded waveform timeline editor for micro-segment-level audio manipulation.
  • Live system telemetry tracking (CPU, RAM, GPU VRAM usage).
  • Targeted multi-speaker diarization — auto-assign unique voice profiles per active speaker.
  • Studio project persistence — save, load, and cache multi-track projects seamlessly via local SQLite.
  • Production SRT/VTT subtitle export packaged alongside the dubbed .mp4 video output.

🔨 Upcoming Features

  • Native Desktop Applications — Dedicated client apps for macOS, Windows, and Linux to entirely bypass CLI execution requirements.
  • Batch Processing — Queue folders full of media to be processed seamlessly overnight.
  • Micro-Tuning Interface — Rapidly fine-tune models to capture micro-expressions with minimal extra reference data.
  • Frame-Perfect Lip-Sync Generation — Align dynamic phoneme synthesis directly to detected on-screen human mouth movement.
  • Universal Voice Plugin System — Decouple the backend to allow bringing your own TTS engines (like XTTS, Bark, etc.).
  • One-Click Deployment — Docker image packages engineered for zero-config GPU passthrough.

📝 Changelog

v1.1.0 — The Cinematic Studio Update

  • The Cinematic Studio Interface: Exhaustively re-engineered the UI to prioritize a high-density, real-estate optimized workflow featuring a dynamic UI zoom scalar (Small, Normal, Max). We minimized dead space and overhauled the widget layout keeping crucial tuning metrics immediately accessible.
  • Multi-Track Timeline: Deeply integrated a multi-layered waveform sequence interface supporting precision audio segment positioning, unmuted live preview playback, localized track timing, and unconstrained draggable positioning manipulation.
  • Persistent Local Projects: Put a complete stop to ephemeral state loss. All workspace metrics are successfully wrapped into Projects logged directly within a native embedded SQLite database. Workflows reliably survive browser shutdowns or server API reboots.
  • AI Cast Diarization: Dropped in an offline Pyannote + WhisperX fusion pipeline evaluating multi-speaker metadata and categorizing overlapping, distinct speakers. Rapidly "cast" clone overrides seamlessly over complex dialogue tracks.
  • Polishing & Asset Control: Cleaned cross-stack filename parsing and exported media rendering via ffmpeg, stabilizing codec dependencies, and deployed a unified custom OmniVoice Studio scalable aesthetic asset system.

Contributions and conceptual ideas are greatly appreciated — open an issue or submit a PR.
S
Description
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
https://voicestudio.sh Readme AGPL-3.0
113 MiB
Languages
Python 51.6%
JavaScript 24.3%
TypeScript 15.9%
Rust 4.8%
CSS 1.7%
Other 1.7%