debpalashandClaude Opus 4.7 d1fd0e5fcb chore: release docs, pin python version, drop stale tarball
- Add docs/RELEASING.md, DESKTOP_RELEASE.md, desktop-build.md for
  release workflow and packaging steps
- Relocate next.md → docs/specs/studio-v1.md (scratch → formal spec)
- Pin Python version via .python-version
- Ignore research/ clones in .gitignore
- Remove stale omnivoice-studio-20260421-1834.tar.gz snapshot

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 17:48:45 +05:30

OmniVoice Logo

OmniVoice Studio

Your Local Cinematic AI Dubbing Studio

FeaturesGetting StartedRoadmapChangelog


OmniVoice Studio Interface Demo
The timeline-based cinematic dubbing and workspace UI.

Local, full-stack voice generation and cinematic dubbing. No API keys. No cloud. Just run it. Built on the open-source OmniVoice 600-language zero-shot diffusion model.

Features

  • 🎬 Video Dubbing — transcribe, translate, re-voice, and mux back into MP4 with selective track export.
  • 🎧 Vocal Isolation — built-in demucs automatically splits speech from music, keeping original background audio perfectly preserved.
  • 🧬 Voice Cloning & Design — Clone specific voices from just a 3-second audio clip, or design completely new studio profiles with tags like female, british accent, excited.
  • Cross-Platform Native Execution — Auto-detects and accelerates inference using Apple Silicon (MPS), NVIDIA (CUDA), AMD (ROCm), or standard CPU.
  • 🔊 Per-Segment Mixing — Fine-grained volume/gain control per dubbed segment (0200%) for broadcast-quality audio balancing.
  • ⌨️ Keyboard-Driven Workflow⌘+Enter to generate, ⌘+S to save, ⌘+Z/⌘+Shift+Z for undo/redo.
  • 📡 Live Model Telemetry — Real-time CPU/RAM/VRAM stats + model warm-up indicator (idle → loading → ready).

🚀 Getting Started

The easiest way to run OmniVoice Studio locally or on a cloud VM is via Docker. Our environment utilizes an optimized pytorch/pytorch configuration which seamlessly enables zero-config GPU passthrough if your host supports it.

git clone https://github.com/debpalash/OmniVoice-Studio.git
cd OmniVoice-Studio

docker compose up --build -d

That's it! Open http://localhost:8000 in your browser.

Tip

Windows/WSL Users: Make sure your NVIDIA drivers are up to date. Docker Desktop automatically passes GPU capabilities to this container! Cloud VMs (AWS, RunPod): The image inherently supports CUDA 12.1. As long as nvidia-container-toolkit is installed on your host, --gpus all binds natively.

Option 2: Local Development Setup

Quickly get OmniVoice Studio running natively on your hardware if you want to develop or modify code. Prerequisites: Ensure ffmpeg is installed on your system. Install standard modern web tooling: Bun and uv.

git clone https://github.com/debpalash/OmniVoice-Studio.git
cd OmniVoice-Studio

# Boot the Backend
uv sync
uv run uvicorn backend.main:app

# Boot the Frontend (in a separate terminal)
bun install
bun run dev

OmniVoice Studio launches exactly two micro-services:

Service Protocol Details
Frontend http://localhost:5173 The real-time React UI — spanning cloning, design, and audio workspace.
Backend http://localhost:8000 The FastAPI server handling model inference, translation pipelines, transcriber tasks.

Note

First run optimization: Model weights (approx. 1.2 GB) automatically download from HuggingFace the first time you execute a generation sequence. Subsequent launches trigger instantly from cache. (Tip: Set HF_TOKEN in your environment for faster, authenticated downloads!)


🗺️ Roadmap

The studio is highly functional today, but we are aggressively expanding. Watch the roadmap to see what's shipping next:

🌟 Completed Milestones

  • Zero-shot voice cloning & complex voice design.
  • Full video cinematic dubbing pipeline (transcribe → translate → synthesize → mux).
  • Vocal isolation utilizing demucs alongside background audio retention.
  • Embedded waveform timeline editor for micro-segment-level audio manipulation.
  • Live system telemetry tracking (CPU, RAM, GPU VRAM usage).
  • Targeted multi-speaker diarization — auto-assign unique voice profiles per active speaker.
  • Studio project persistence — save, load, and cache multi-track projects seamlessly via local SQLite.
  • Production SRT/VTT subtitle export packaged alongside the dubbed .mp4 video output.
  • Selective track export — choose exactly which language tracks (Original, DE, ES, etc.) to include in final MP4.
  • Per-segment volume/gain control with real-time mixing (0200%).
  • Undo/redo system for all segment edits with 50-action history depth.
  • Keyboard shortcuts: ⌘+Enter generate, ⌘+S save, ⌘+Z/⌘+Shift+Z undo/redo.
  • Drag-and-drop file uploads for both video and clone audio sources.
  • Model warm-up indicator with live status pill (idle/loading/ready).
  • Confirmation dialogs for all destructive actions (delete project/history/profile).
  • UI preferences persistence (sidebar state, zoom, active tab) across sessions.
  • Polished glassmorphism design system with micro-animations, focus rings, and custom scrollbars.

🔨 Upcoming Features

  • Real Speaker Diarization — ML-based diarization via pyannote.audio for true multi-speaker identification.
  • A/B Voice Comparison — Side-by-side voice audition for casting decisions.
  • Scene-Aware Dubbing — FFmpeg scene detection to auto-split segments at visual cuts.
  • Lip-Sync Scoring — Analyze dubbed audio duration against original speaker timing with color-coded badges.
  • Batch Processing — Centralized async task queue ensuring sequential GPU execution with reconnectable SSE streams.
  • Advanced Export Suite — VTT subtitles, per-segment WAV ZIP, compressed MP3, and stem export (vocals + background separate).
  • Streaming TTS — Chunked WAV streaming with progressive download and auto-playback.
  • Native Desktop Applications — Dedicated client apps for macOS, Windows, and Linux.
  • One-Click Deployment — Docker image packages engineered for zero-config GPU passthrough.

📝 Changelog

v1.2.0 — The Production Polish Update

  • Selective Track Export: Choose exactly which audio tracks to include in the final MP4. Uncheck Original, keep only German — get a single-track export. Full per-track checkbox UI with dynamic FFmpeg stream index remapping.
  • Undo/Redo System: Full ⌘+Z / ⌘+Shift+Z undo/redo for all segment edits (text, voice, volume, delete). 50-action deep history stack.
  • Per-Segment Volume Control: Inline gain slider (0200%) per segment row in the dub table. Backend applies gain during audio assembly with safe clamping.
  • Keyboard Shortcuts: ⌘+Enter to generate, ⌘+S to save project. Browser default overrides prevented.
  • Model Status Indicator: Live status pill in the header showing model warm-up state (idle → loading → ready). New /model/status backend endpoint.
  • Drag-and-Drop Everywhere: Video upload already supported drop — now clone audio upload does too, with pink highlight on hover.
  • Confirmation Dialogs: All destructive actions (delete project, profile, history item, clear all history) now require confirmation.
  • Session Persistence: Sidebar collapsed state, active tab, and zoom level now persist across browser sessions via localStorage.
  • CSS Design System Overhaul: Anti-aliased text, input focus glow rings, button hover shimmer, progress bar shimmer animation, fade-in on history items, selection color branding, Firefox scrollbar support, tabular-nums for timestamp columns.
  • AudioContext Pooling: playPing() synthesis notification reuses a single AudioContext instead of creating one per call (browsers cap at ~6).

v1.1.0 — The Cinematic Studio Update

  • The Cinematic Studio Interface: Exhaustively re-engineered the UI to prioritize a high-density, real-estate optimized workflow featuring a dynamic UI zoom scalar (Small, Normal, Max). We minimized dead space and overhauled the widget layout keeping crucial tuning metrics immediately accessible.
  • Multi-Track Timeline: Deeply integrated a multi-layered waveform sequence interface supporting precision audio segment positioning, unmuted live preview playback, localized track timing, and unconstrained draggable positioning manipulation.
  • Persistent Local Projects: Put a complete stop to ephemeral state loss. All workspace metrics are successfully wrapped into Projects logged directly within a native embedded SQLite database. Workflows reliably survive browser shutdowns or server API reboots.
  • AI Cast Diarization: Dropped in an offline Pyannote + WhisperX fusion pipeline evaluating multi-speaker metadata and categorizing overlapping, distinct speakers. Rapidly "cast" clone overrides seamlessly over complex dialogue tracks.
  • Polishing & Asset Control: Cleaned cross-stack filename parsing and exported media rendering via ffmpeg, stabilizing codec dependencies, and deployed a unified custom OmniVoice Studio scalable aesthetic asset system.

Star History


Contributions and conceptual ideas are greatly appreciated — open an issue or submit a PR.
S
Description
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
https://voicestudio.sh Readme AGPL-3.0
113 MiB
Languages
Python 51.6%
JavaScript 24.3%
TypeScript 15.9%
Rust 4.8%
CSS 1.7%
Other 1.7%