diff --git a/CHANGELOG.md b/CHANGELOG.md index c9eaecf9..25f87054 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,14 @@ the frozen-backend fallback mirror it for their toolchains. **Highlights** +- The README is shorter, with a new Electron UI tour and refreshed screenshots (#2129) + +- Support pages feature cleaner donation cards, with a workspace support shortcut and sponsor footer with hover cards and email inquiries (#2129) + +- Integrations has a dedicated sidebar workspace with featured sponsors, searchable AI providers, and smooth sponsor-strip scrolling (#2129) + +- Integrations now covers 100+ automation, communications, MCP, agent, developer, data, and productivity tools with config-driven detail pages (#2129) + - Electron now ships as a complete cross-platform VoiceStudio desktop app with local-first cloning, production workspaces, model packs, repair agents, native integrations, updates, parity checks, and the shared backend contracts required by those workflows (#1823) - The Model Catalogue is one page: what you use now on top, then each family's engines and weights (#2013) diff --git a/README.md b/README.md index cb8d6cdf..b1ad3a4b 100644 --- a/README.md +++ b/README.md @@ -1,502 +1,81 @@
- -

NOTE: Electron Rewrite Ongoing: Please dont't create desktop app related issues and pr

- -

VoiceStudio logo

+ VoiceStudio

VoiceStudio

+

Your voices. Your stories. Your machine.

+

Clone voices, dub videos, dictate, and create audiobooks with local AI.

- VoiceStudio ranking on Trendshift -

-

Previously OmniVoice-Studio

-

Clone voices, dub video, dictate, and produce long-form audio on your own hardware.

-

16 TTS engines · 11 ASR engines · 646-language catalogue · macOS, Windows, Linux, and Docker

-

No account, API key, subscription, or usage meter for the local workflow.

- -

- Install · - Features · - Compare · - Requirements · - Hardware · - Engines · - Architecture · - API · + Download · + Get started · Docs · - FAQ · - 简体中文 + Discord · + 简体中文

-

- CI status - GitHub stars - Total downloads - Latest release - AGPL-3.0 license - Discord community -

- -

- Download VoiceStudio + CI + Latest release + AGPL-3.0

-
- Switching TTS engines from the VoiceStudio status bar -
+![A tour of the Electron app: voice cloning, voice design, dubbing, and model management](docs/media/electron/voicestudio.gif) -> [!WARNING] -> **Active beta.** Use the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) for stable work. `main` contains the newest fixes and may change between releases. Report problems through [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues). +

The new Electron desktop UI, captured from this branch with the bundled demo voice. Release builds may look different.

-## At a glance +## Create with VoiceStudio -| | VoiceStudio | -|---|---| -| **Workflows** | Voice cloning and design, video dubbing, dictation, stories, audiobooks, batch generation | -| **Language catalogue** | 646 TTS languages; actual coverage and quality depend on the selected engine | -| **Engines** | 16 TTS · 11 ASR · switch in Model Catalogue or with Ctrl/Cmd+E | -| **Platforms** | macOS 13.3+ on Apple Silicon · Windows 10/11 x64 · Linux x86_64 with glibc 2.39+ | -| **Compute** | CUDA · Apple Silicon MPS/MLX · ROCm on Linux · CPU · optional remote workers | -| **Interfaces** | Desktop app · local REST/SSE/WebSocket API · OpenAI-compatible audio API · MCP Server | -| **Storage** | Voices, projects, settings, and outputs stay on the machine by default | -| **License** | AGPL-3.0 application; downloaded models keep their upstream terms | +- **Clone & design voices** — use a reference recording or describe the voice you imagine. +- **Dub video** — transcribe, translate, assign speakers, and edit timed speech. +- **Dictate anywhere** — record, transcribe, and copy text with a floating recording widget. +- **Tell longer stories** — create multi-voice scripts, audiobooks, and batch jobs. +- **Choose your models** — manage speech and transcription engines, languages, and compute devices. -The Voice workspace starts with three tabs: **From audio** for cloning, **By design** for creating a voice, and **Convert** for speech-to-speech conversion. Each tab displays its own workflow, with Synthesize Audio or Convert pinned below the scrolling form. The top-bar **Engines** panel combines engine selection, loaded models, and unload/flush controls; Ctrl/Cmd+E opens it. The searchable language picker shares Dubbing’s flags and language list layout, selects one output language, and retains Auto and the full cloning catalogue. Language options flow into multiple columns when space allows. Expand **Workspaces** in the sidebar to reveal navigation labels; Escape collapses it. +Start with **VoiceStudio** (default, powered by k2-fsa/OmniVoice), or choose another engine. -Dubbing starts with file upload or URL import and nearby language choices. Its **Projects** panel lists previous dubs so they can be reopened by clicking anywhere on a card; action buttons operate independently. Advanced import options include captions and optional YouTube sign-in. Dubbing places playback controls over the video with background blur and combines the waveform and timed transcript in one compact editing surface. Drag the zoomed waveform left or right to pan; click to seek. Translation language and ISO-code controls stay synchronized; Auto clears any previous language code and dialect. Transcript items group editable text, timing and status, and voice controls into three readable rows that wrap with the panel width. Output Options stays compact with the active settings shown in its summary; expand it to change output, timing, or voice matching. Transcript, glossary, and paste controls share a toolbar above the segment editor. Project details, workflow steps, and Generate/Verify/Export actions use an unfilled header. +Local workflows run on your hardware. Remote services are optional; usage analytics requires consent. -The Audiobook Script editor fills the available workspace beneath its markup toolbar; Voices and Book settings stay in their own tabs. + + + + + + +
Electron voice cloning workspace with the bundled demo voiceElectron video dubbing workspace
Voice cloningVideo dubbing
-Output settings use aligned rows; review status appears before the collapsible transcript and glossary. Glossary terms have labelled entry fields and an explicit edit action. Launchpad arranges recent files and saved voices side by side when space allows, with responsive card grids and visible Open actions. +## Get started -The casting board shows icon-based voice cards and searchable selectors for each speaker. Drag a card onto a speaker or choose a voice from that speaker’s menu. +Download from [Releases](https://github.com/debpalash/VoiceStudio/releases/latest), then follow your platform guide: - +**[macOS](docs/install/macos.md) · [Windows](docs/install/windows.md) · [Linux](docs/install/linux.md) · [Docker](docs/install/docker.md)** -## Install +Open **Voice cloning**, choose a voice or add a clean reference recording, enter your text, and generate. Install the required model when prompted. Hardware needs vary by engine; see [performance](docs/performance.md). -Download a package from the [latest release](https://github.com/debpalash/VoiceStudio/releases/latest), then follow the platform guide. - -| Platform | Package | Guide | -|---|---|---| -| macOS 13.3+ | Apple Silicon DMG | [Install on macOS](docs/install/macos.md) | -| Windows 10/11 | x64 MSI; choose the current-user build when listed to install without admin access | [Install on Windows](docs/install/windows.md#install-pre-built-msi) | -| Linux | AppImage, x86_64 with glibc 2.39+ | [Install on Linux](docs/install/linux.md) | -| Docker | Linux/AMD64 images; CUDA, ROCm, CPU, and worker-only GPU profiles | [Run with Docker](docs/install/docker.md) | - -First launch creates a managed Python environment and downloads the default model. Later launches reuse both. - -> [!NOTE] -> On macOS, first launch needs a one-time right-click, then **Open** approval. Intel Macs cannot run the local Python backend; use a [remote backend](docs/install/macos.md) instead. - -### Quick Docker run - -The published images are **`linux/amd64` only**. On Apple Silicon, use the -[native macOS app](docs/install/macos.md) for GPU acceleration. ARM64 hosts -should read the [architecture requirements](docs/install/docker.md#architecture) -before pulling an image. - -```bash -docker run -d -p 127.0.0.1:3900:3900 -v omnivoice-data:/app/omnivoice_data --name voicestudio palashdeb/omnivoice-studio:stable -``` - -### First voice - -1. Launch VoiceStudio and open **Voice Cloning**. -2. Add a clean voice sample. Three seconds works; 5 to 15 seconds usually gives a better prompt. -3. Enter text, choose a language, then select **Generate**. - -> [!TIP] -> **Try without installing:** Run VoiceStudio in the cloud via the [Google Colab notebook](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb). Explore audio quality comparisons in [benchmarks](docs/benchmarks.md) and prompt design tips in [expressive speech](docs/expressive-speech.md). - -### Audio samples - -Listen to sample outputs produced locally with VoiceStudio: - -| Workflow | Prompt / Reference Audio | Generated Audio | -|---|---|---| -| **Voice Cloning** | [demo_voice.wav](backend/assets/samples/demo_voice.wav) | [demo_clone_output.wav](backend/assets/samples/demo_clone_output.wav) | -| **Voice Design** (US News Anchor) | *"Clear, authoritative American broadcast tone"* | [demo_voice_design_us_news_anchor.wav](backend/assets/samples/voice_design/demo_voice_design_us_news_anchor.wav) | -| **Voice Design** (UK Audiobook) | *"Warm, expressive British storytelling voice"* | [demo_voice_design_audiobook_uk_narrator.wav](backend/assets/samples/voice_design/demo_voice_design_audiobook_uk_narrator.wav) | -| **Video Dubbing** (Multilingual) | [source.src.wav](backend/assets/samples/demo/dubbing/source.src.wav) | [Spanish](backend/assets/samples/demo/dubbing/dubbed_es.src.wav) · [French](backend/assets/samples/demo/dubbing/dubbed_fr.src.wav) · [Japanese](backend/assets/samples/demo/dubbing/dubbed_ja.src.wav) · [Chinese](backend/assets/samples/demo/dubbing/dubbed_zh.src.wav) | - -### Run from source - -Install the [development prerequisites](.github/CONTRIBUTING.md#development-setup) (Node 20+/Bun and Python 3.11+), then: +**Run the Electron preview from source:** ```bash git clone https://github.com/debpalash/VoiceStudio.git cd VoiceStudio bun install -bun run desktop +cd electron +bun run dev ``` -The desktop launcher configures Python dependencies on first run via `uv` automatically. Use `bun run dev` for the browser UI. See [Contributing](.github/CONTRIBUTING.md) for services, tests, and platform packages. - -### If setup fails - -- Run **Settings → About → Run self-check** or `uv run python backend/main.py --diagnose --deep`. -- Check [install troubleshooting](docs/install/troubleshooting.md). -- Save a scrubbed diagnostic bundle from the app when opening an issue. -- For slow generation, compare [measured benchmarks](docs/benchmarks.md) and [performance settings](docs/performance.md). - - - -## Features - -| Area | Included | -|---|---| -| **Voice Cloning** | Zero-shot synthesis from a short reference clip ([guide](docs/engines/README.md)) | -| **Voice Design** | Create a voice from age, accent, pitch, style, and delivery instructions ([expressive speech](docs/expressive-speech.md)) | -| **Video Dubbing** | Transcribe, translate, preserve speakers, synthesize, and export video; compact translation settings include track selection, and completed dubs flag timing issues for review ([export guide](docs/dubbing/export.md)) | -| **Stories and audiobooks** | Multi-voice scripts · EPUB/PDF import · chapter rendering · `.m4b` export | -| **[Dictation Widget](docs/features/dictation.md)** | System-wide shortcut, live transcription, optional local-LLM cleanup | -| **Vocal Isolation** | Demucs speech/background separation | -| **Speaker Diarization** | Pyannote and WhisperX speaker assignment ([guide](docs/features/diarization.md)) | -| **Batch Queue** | Queue large sets of audio and video jobs with per-job progress, or watch a local folder for new videos | -| **Model Catalogue** | Install, remove, select, and route TTS, ASR, and LLM models ([catalogue](docs/engines/README.md)) | -| **Remote Model Downloads** | Install models on enrolled remote workers with live progress ([guide](docs/downloading-models.md)) | -| **GPU Auto-Detect** | CUDA, MPS, ROCm, and CPU routing with per-engine checks ([performance](docs/performance.md)) | -| **AI Watermark** | AudioSeal embedding and detection | -| **MCP Server** | Synthesis and transcription tools for MCP clients ([guide](docs/mcp.md)) | -| **Diagnostics** | Self-checks, error journal, logs, and scrubbed support bundles ([troubleshooting](docs/install/troubleshooting.md)) | -| **Local-first** | Core creation stays local; network-backed features are explicit opt-ins | -| **Extensible** | Registry-based TTS, ASR, and plugin interfaces ([acceptance](docs/engine-acceptance.md)) | - - - - - - - - - - -
VoiceStudio Model CatalogueSaving a gallery voice as a local profile
Model Catalogue: engine, device, and install stateGallery: save a shared voice as a local profile
- - - -## Comparison - -VoiceStudio trades managed cloud compute for local control. This is the practical difference: - -| | **VoiceStudio** | **Typical hosted voice service** | -|---|---|---| -| **Best fit** | Private, offline, self-hosted, or high-volume work | Fast setup without local model management | -| **Data path** | Local by default; remote features are opt-in | Audio and text are processed by the provider | -| **Cost model** | Free software; you supply the hardware | Subscription, credits, or metered API use | -| **Setup** | Install the app and model weights | Create an account and use the web app or API | -| **Performance** | Depends on your engine and hardware | Provider manages compute and scaling | -| **Offline use** | Yes, after required models are installed | Usually requires a network connection | -| **Customization** | Source, engines, models, API, and routing are open | Limited to provider options | -| **Maintenance** | You manage updates, disk, and compute | Provider manages infrastructure | - - - -## Requirements - -Requirements vary by engine. These values cover the default local workflow. - -| | **Minimum** | **Recommended** | -|---|---|---| -| **OS** | Windows 10 x64 · macOS 13.3 Apple Silicon · Linux x86_64 with glibc 2.39+ | Current supported OS release | -| **RAM** | 8 GB | 16 GB+ | -| **Disk** | 10 GB free | 20 GB+ SSD | -| **GPU** | Optional; CPU mode is supported | NVIDIA CUDA or Apple Silicon | -| **VRAM** | 4 GB when using a GPU | 8 GB+; large optional engines need more | -| **Python from source** | 3.11+ | 3.11 or 3.12 | - -ROCm is Linux-only and opt-in. Windows AMD/Ryzen AI uses CPU. Systems with limited VRAM offload work to CPU when required. See [performance](docs/performance.md), [benchmarks](docs/benchmarks.md), and [engine disk usage](docs/engines/disk-usage.md). - - - -### Recommended stack by hardware - -| Hardware | Recommended TTS | Recommended ASR | Why | -|---|---|---|---| -| **Apple Silicon (M1–M4)** | [MLX-Audio](docs/engines/mlx-audio.md) · [OmniVoice](docs/engines/omnivoice.md) (MPS) | [MLX Whisper](docs/engines/mlx-whisper.md) · [Parakeet MLX](docs/engines/parakeet-mlx.md) | Native unified memory, lowest latency on macOS | -| **NVIDIA GPU (8 GB+ VRAM)** | [OmniVoice](docs/engines/omnivoice.md) · [CosyVoice 3](docs/engines/cosyvoice.md) | [WhisperX](docs/engines/whisperx.md) | High-fidelity zero-shot cloning, word timestamps, diarization | -| **Low VRAM / CPU-only** | [PocketTTS](docs/engines/pockettts.md) · [Sherpa-ONNX](docs/engines/sherpa-onnx.md) · [KittenTTS](docs/engines/kittentts.md) | [Moonshine](docs/engines/moonshine.md) · [Faster-Whisper](docs/engines/faster-whisper.md) (`int8`) | Low memory footprint, optimized CPU inference | - - - -## Engines - -Engine support is capability-specific. Check cloning, language, platform, memory, and license before choosing one. Full setup guides: [docs/engines](docs/engines/README.md). - - - -### Text to speech - -| Engine | Languages | Clone | Instruct | Linux | macOS ARM | Windows | License | -|---|:---:|:---:|:---:|:---:|:---:|:---:|---| -| [**VoiceStudio** (default, powered by k2-fsa/OmniVoice)](docs/engines/omnivoice.md) | 600+ | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0 code, CC-BY-NC weights](https://huggingface.co/k2-fsa/OmniVoice#license)³ | -| [**CosyVoice 3**](docs/engines/cosyvoice.md) | 9 + 18 dialects | Yes | Yes | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 | -| [**GPT-SoVITS**](docs/engines/gpt-sovits.md) | 5 | Yes | No | CUDA/CPU | No | CUDA/CPU | MIT | -| [**VoxCPM2**](docs/engines/voxcpm2.md) | 30 | Yes | Yes | CUDA/CPU | MPS | CUDA/CPU | Apache-2.0 | -| [**MOSS-TTS-Nano**](docs/engines/moss-tts-nano.md) | 20 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 | -| [**KittenTTS**](docs/engines/kittentts.md) | English | No | No | CPU | CPU | CPU | MIT | -| [**MLX-Audio**](docs/engines/mlx-audio.md) | Model-dependent | Varies | Varies | No | MLX | No | Varies | -| [**Sherpa-ONNX**](docs/engines/sherpa-onnx.md) | 20+ | No | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 | -| [**IndexTTS 2.5** ⚡](docs/engines/indextts.md) | ZH · EN · JA · ES · AR | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Bilibili model license¹ | -| [**OmniVoice GGUF** ⚡](docs/engines/omnivoice-gguf.md) | 600+ | Yes | Yes | CUDA/CPU | MPS/CPU | CUDA/CPU | [AGPL-3.0](LICENSE) app · [review the derivative model terms](https://huggingface.co/Serveurperso/OmniVoice-GGUF#license)³ | -| [**OmniVoice (subprocess; opt-in off MPS)** ⚡](docs/engines/omnivoice-subprocess.md) | 600+ | Yes | Yes | CUDA/CPU | MPS via default OmniVoice | CUDA/CPU | [AGPL-3.0](LICENSE) app · [Apache-2.0 code, CC-BY-NC weights](https://huggingface.co/k2-fsa/OmniVoice#license)³ | -| [**PocketTTS** ⚡](docs/engines/pockettts.md) | EN · FR · DE · PT · IT · ES | Yes | No | CPU | CPU | CPU | CC-BY-4.0, gated² | -| [**Supertonic 3** ⚡](docs/engines/supertonic3.md) | 31 | No | No | CPU | CPU | CPU | OpenRAIL-M | -| [**MOSS-TTS-v1.5** ⚡](docs/engines/moss-tts-v15.md) | 31 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 | -| [**dots.tts** ⚡](docs/engines/dots-tts.md) | 24 | Yes | No | CUDA/CPU | CPU | No | Apache-2.0 | -| [**Confucius4-TTS** ⚡](docs/engines/confucius4-tts.md) | 14 | Yes | No | CUDA/CPU | CPU | CUDA/CPU | Apache-2.0 | - -⚡ Installed or registered on demand. - -¹ IndexTTS 2.5 requires a separate written Bilibili license above 100 million monthly active users or RMB 1 billion annual revenue. Review the [model license](https://huggingface.co/IndexTeam/IndexTTS-2.5/blob/main/LICENSE). - -² PocketTTS shows its gated-access and CC-BY-4.0 terms before first use. - -³ The OmniVoice snapshot also includes an audio tokenizer under separate [Boson Higgs Audio 2 and Meta Llama community terms](https://huggingface.co/k2-fsa/OmniVoice/blob/main/audio_tokenizer/LICENSE). VoiceStudio's application license does not replace model or tokenizer terms. - -Clone-less engines cannot preserve a reference speaker in dubbing or pinned-voice batch jobs. VoiceStudio rejects those jobs instead of silently changing engines. Heavy engines have separate memory and platform limits; check their engine guide first. - - - -### Speech to text - -| Engine | ID | Languages | Best fit | -|---|---|:---:|---| -| [**WhisperX** (default)](docs/engines/whisperx.md) | `whisperx` | ~100 | Dubbing, subtitles, word-level timing | -| [**Faster-Whisper**](docs/engines/faster-whisper.md) | `faster-whisper` | ~100 | General cross-platform transcription | -| [**Faster-Whisper (isolated)**](docs/engines/faster-whisper-isolated.md) | `faster-whisper-isolated` | ~100 | Crash-isolated batch transcription | -| [**MLX Whisper**](docs/engines/mlx-whisper.md) | `mlx-whisper` | ~100 | Apple Silicon | -| [**PyTorch Whisper**](docs/engines/pytorch-whisper.md) | `pytorch-whisper` | ~100 | CUDA, MPS, and CPU fallback | -| [**Parakeet TDT**](docs/engines/nemo-parakeet.md) | `nemo-parakeet` | English + 25 EU | Fast CPU/CUDA transcription | -| [**Parakeet TDT v3 (MLX)**](docs/engines/parakeet-mlx.md) | `parakeet-mlx` | 25 EU | Apple Silicon dictation and word timestamps | -| [**Moonshine**](docs/engines/moonshine.md) | `moonshine` | English | Low-power, low-latency ONNX | -| [**FunASR**](docs/engines/funasr.md) | `funasr` | 50+ | VAD and inline diarization | -| [**sherpa-onnx** (live dictation)](docs/engines/sherpa-onnx-asr.md) | `sherpa-onnx-asr` | Model-dependent | Streaming CPU dictation | -| [**OpenAI-compatible** ⚠️ configured server](docs/engines/openai-compatible-asr.md) | `openai-compat-asr` | Server-dependent | Local gigastt/Qwen3-ASR or a remote endpoint; audio goes only to that server | - -WhisperX and Faster-Whisper retry with `int8` when efficient `float16` is unavailable. Pin `ASR_COMPUTE_TYPE=int8` or `float32` only if automatic selection still fails. - - - -## Architecture - -```text -Tauri v2 desktop shell (Rust) - │ IPC -React + Vite UI - │ HTTP · SSE · WebSocket on localhost:3900 -FastAPI backend - ├── TTS / ASR engine registries - ├── dubbing / audio / long-form pipelines - ├── OpenAI-compatible API and MCP server - └── SQLite + Alembic → omnivoice_data/ -``` - -| Layer | Path | Responsibility | -|---|---|---| -| Desktop shell | `frontend/src-tauri/` | Window lifecycle, tray, shortcuts, updater, sidecar bootstrap | -| Frontend | `frontend/src/` | React UI, Zustand state, API and event clients, i18n | -| API | `backend/api/` | REST routes, schemas, auth boundaries, streaming | -| Core services | `backend/services/` | Generation, dubbing, audio processing, persistence | -| Engines | `backend/engines/` | Isolated and optional engine adapters | -| Worker system | `backend/worker/` | Authenticated remote compute and job transport | -| Data | `omnivoice_data/` | Projects, voices, settings, logs, and SQLite state | -| Delivery | `scripts/`, `deploy/`, `.github/workflows/` | Development, packaging, containers, releases, CI | - -### Network boundary - -- The desktop talks to a loopback-only backend on `localhost:3900`. -- Loopback API calls need no server key. Remote access requires a share PIN or API key. -- Remote workers and OpenAI-compatible ASR are opt-in. Loopback ASR may use HTTP and keeps audio on the machine; non-loopback endpoints require HTTPS, and redirects are not followed. -- Analytics is off until consent. If enabled, it sends allowlisted, content-free usage metadata. It never sends text, audio, file names, or projects. - - - -## Local speech platform and OpenAI-compatible API - -Point an OpenAI-compatible audio client at the local backend: - -```diff -- base_url="https://api.openai.com/v1" -+ base_url="http://localhost:3900/v1" -``` - -| Endpoint | Purpose | -|---|---| -| `POST /v1/audio/speech` | TTS to `mp3`, `opus`, `aac`, `flac`, `wav`, or `pcm`; select a profile with `voice` and an engine with `model` | -| `POST /v1/audio/transcriptions` | STT to `json`, `text`, `verbose_json`, `srt`, or `vtt` | -| `WS /v1/audio/transcriptions/stream` | Live PCM/WebM transcription with partial, utterance, and session-final events | -| `GET /.well-known/voicestudio-speech` | Discover HTTP, WebSocket, MCP, and native dictation-control transports | -| `GET /v1/audio/voices` | List local voice profiles and engines | - -```python -from openai import OpenAI - -client = OpenAI(base_url="http://localhost:3900/v1", api_key="local") - -with client.audio.speech.with_streaming_response.create( - model="tts-1", - voice="", - input="Made on my own hardware.", - response_format="wav", -) as response: - response.stream_to_file("speech.wav") -``` - -```bash -# Quick test via cURL -curl http://localhost:3900/v1/audio/speech \ - -H "Content-Type: application/json" \ - -d '{"model": "tts-1", "input": "Made on my own hardware.", "voice": "default", "response_format": "wav"}' \ - --output speech.wav -``` - -The bundled Rust control sidecar lets Herdr, coding agents, VS Code, desktop apps, -and TUIs trigger the system-wide dictation flow or reuse its native text -insertion. See the [speech platform guide](docs/speech-platform.md). The full API -reference is in **Settings → OpenAPI Reference**. For LAN, Tailscale, or proxy -access, read [API authentication](docs/api-auth.md) before exposing the backend. - -### Agent skills - -Install the VoiceStudio skills for Claude Code, Codex, Cursor, and other [skills.sh](https://skills.sh)-compatible agents: - -```bash -npx skills add debpalash/VoiceStudio -``` - -- `omnivoice`: synthesize speech and transcribe audio through local VoiceStudio. -- `oss-maintainer`: the repository's open-source maintenance workflow. - -### Model Context Protocol (MCP) - -VoiceStudio mounts an MCP server at `http://localhost:3900/mcp` for Claude Desktop, Cursor, and AI agents: - -```json -{ - "mcpServers": { - "voicestudio": { - "url": "http://localhost:3900/mcp" - } - } -} -``` - -For clients requiring stdio transport, use the bundled local shim (`docs/mcp.json`): - -```json -{ - "mcpServers": { - "voicestudio": { - "command": "python", - "args": ["-m", "backend.mcp_shim"], - "cwd": "/path/to/VoiceStudio" - } - } -} -``` - -See the [MCP guide](docs/mcp.md) for tools (`generate_speech`, `clone_voice`, `transcribe`), file streaming modes, and client bindings. - -### Google Colab - -[![Open in Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb) - -The [notebook](notebooks/OmniVoice_Studio_Colab.ipynb) runs the app and web UI on a Colab GPU. Colab is remote compute, so uploaded audio and project data do not remain local to your machine. - - +See [Electron setup](electron/README.md) for prerequisites and backend configuration. VoiceStudio is in active development; report bugs through [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues). ## Documentation -| Need | Read | +| Need | Start here | |---|---| -| Install | [macOS](docs/install/macos.md) · [Windows](docs/install/windows.md) · [Linux](docs/install/linux.md) · [Docker](docs/install/docker.md) | -| Fix setup | [Troubleshooting](docs/install/troubleshooting.md) · [model downloads](docs/downloading-models.md) · [Hugging Face token](docs/setup/huggingface-token.md) | -| Choose an engine | [Engine guides](docs/engines/README.md) · [benchmarks](docs/benchmarks.md) · [expressive speech](docs/expressive-speech.md) | -| Tune hardware | [Performance](docs/performance.md) · [remote workers](docs/remote-workers.md) | -| Build integrations | [Speech platform](docs/speech-platform.md) · [Private production API](docs/production-private-api.md) · [API auth](docs/api-auth.md) · [MCP](docs/mcp.md) · [examples](examples/README.md) | -| Build VoiceStudio | [Contributing](.github/CONTRIBUTING.md) · [engine acceptance](docs/engine-acceptance.md) | -| Track changes | [Changelog](CHANGELOG.md) · [roadmap](docs/ROADMAP.md) · [latest release](https://github.com/debpalash/VoiceStudio/releases/latest) | -| Remove everything | [Uninstall guide](docs/install/uninstall.md) | +| Setup help | [Troubleshooting](docs/install/troubleshooting.md) · [Model downloads](docs/downloading-models.md) | +| Models & audio quality | [Engine guides](docs/engines/README.md) · [Benchmarks](docs/benchmarks.md) | +| Integrations | [Local API](docs/speech-platform.md) · [MCP](docs/mcp.md) · [Examples](examples/README.md) | +| Development | [Contributing](.github/CONTRIBUTING.md) · [Electron](electron/README.md) · [Changelog](CHANGELOG.md) | - +Agent skills: `npx skills add debpalash/VoiceStudio` -## FAQ +## Support VoiceStudio -
-Does it work on Apple Silicon and Intel Macs? +[Ko-fi](https://ko-fi.com/debpalash) · [PayPal](https://paypal.me/palashCoder) · [Sponsor the project](SPONSORS.md) · [Partnerships](mailto:partner@voicestudio.sh) -Apple Silicon is supported with MPS and MLX options. Intel Macs cannot run the local backend because current PyTorch wheels are unavailable; they can connect to a remote backend. See [macOS installation](docs/install/macos.md). -
+## License & responsible use -
-How much VRAM do I need? - -A GPU is optional. Use 4 GB VRAM as the minimum for accelerated work and 8 GB+ for the default multi-stage workflow. Large optional engines can require 12 to 16 GB or more. Check the [benchmarks](docs/benchmarks.md) and engine guide. -
- -
-Why does a longer reference clip not always improve the clone? - -Cloning is zero-shot: the clip is a prompt, not training data. Use 5 to 15 seconds of one speaker, close to the microphone, without music, noise, or reverb. Match the tone and pace you want in the output. For training, see [data preparation](docs/data_preparation.md) and [training](docs/training.md). -
- -
-Can I use generated audio commercially? - -VoiceStudio's application license does not restrict generated audio, but it does not grant rights under a model's separate terms. The default OmniVoice repository labels its pretrained weights CC-BY-NC and includes a tokenizer under separate community terms. Review the selected model terms before commercial use. -
- -
-Does VoiceStudio collect data? - -Not unless you opt in. Analytics is off by default and skipping consent keeps it off. When enabled, the app sends allowlisted, content-free usage metadata. Text, audio, file names, voices, and projects are excluded. Change this at **Settings → Privacy**. -
- -
-How do I remove VoiceStudio and its data? - -Use `scripts/uninstall.sh` on macOS/Linux or `scripts\uninstall.ps1` on Windows. Both show a dry run before deletion. See the [uninstall guide](docs/install/uninstall.md) for every path. -
- -## Community and contributing - -- [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues) for reproducible bugs and feature requests. -- [Discord](https://discord.gg/bzQavDfVV9) for setup help and project discussion. -- [Good first issues](https://github.com/debpalash/VoiceStudio/labels/good%20first%20issue) for a scoped starting point. -- [Contributing guide](.github/CONTRIBUTING.md) for setup, tests, and pull requests. - -

- - Star History Chart - -

- -## Support development - -VoiceStudio is free and has no paid tier. Donations fund development and infrastructure. - -[Ko-fi](https://ko-fi.com/debpalash) · [PayPal](https://paypal.me/palashCoder) · [Sponsorship details](SPONSORS.md) - -## Responsible use and safety - -VoiceStudio enables zero-shot voice cloning and speech generation on personal hardware. Please use it responsibly: -- **Consent:** Only clone or synthesize voices with explicit permission from the speaker. -- **Audio provenance:** VoiceStudio integrates [AudioSeal](https://github.com/facebookresearch/audioseal) imperceptible watermarking by default to detect and identify synthetic speech without altering sound quality. -- **Local privacy:** For the default local workflow, audio recordings, transcripts, voices, and projects remain strictly on your local disk; data leaves your device only when you explicitly configure remote workers or external ASR endpoints. - -## License - -VoiceStudio is licensed under [AGPL-3.0](LICENSE). You may run it, modify it, and use it internally. The application license itself does not restrict selling generated audio, but downloaded model and tokenizer terms may. If you modify VoiceStudio and provide that modified version as a network service, AGPL requires you to offer the corresponding source under the same license. A commercial license for VoiceStudio-owned code is available for proprietary embedding; it does not relicense third-party models. Contact **VoiceStudio@palash.dev**. See [LICENSE-NOTICE.md](LICENSE-NOTICE.md) for the plain-language scope. - -Optional engines and downloaded models retain their own licenses. The bundled `omnivoice/` Python code is Apache-2.0 upstream; the default downloaded weights and audio tokenizer use separate terms. - -## Acknowledgments - -VoiceStudio builds on [OmniVoice](https://github.com/k2-fsa/OmniVoice), [WhisperX](https://github.com/m-bain/whisperX), [Demucs](https://github.com/facebookresearch/demucs), [Pyannote](https://github.com/pyannote/pyannote-audio), [CTranslate2](https://github.com/OpenNMT/CTranslate2), [AudioSeal](https://github.com/facebookresearch/audioseal), [Tauri](https://tauri.app), [Supertonic](https://huggingface.co/Supertone/supertonic-3), [Sherpa-ONNX](https://github.com/k2-fsa/sherpa-onnx), [GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS), and [PocketTTS](https://kyutai.org). - -
- Download VoiceStudio · - Star the project · - Join Discord -
+[AGPL-3.0](LICENSE). Models have their own licenses; review them before commercial use. Clone voices only with permission. See [license details](LICENSE-NOTICE.md). diff --git a/README_CN.md b/README_CN.md index 0c998c22..4e4b331d 100644 --- a/README_CN.md +++ b/README_CN.md @@ -1,697 +1,74 @@ -*本文档是 [README.md](README.md) 的简体中文翻译;若与英文版有出入,以英文版为准。* -
- VoiceStudio 徽标 + VoiceStudio

VoiceStudio

-

原名 OmniVoice-Studio

-

创造声音,讲述故事,文件始终属于你。♡

-

在一个开源桌面工作室里完成克隆、设计、配音、听写和有声书制作。
默认本地优先。没有订阅,也没有用量计费;联网服务始终由你主动选择。

- +

你的声音,你的故事,在你的电脑上创作。

+

使用本地 AI 克隆声音、翻译配音、语音听写和制作有声书。

- 快速开始 · - 功能 · - 为什么选择 VoiceStudio · - 引擎 · - API · - 捐赠 · - 参与贡献 · + 下载 · + 开始使用 · + 文档 · Discord · - English -

- -

- CI 状态 - Star 数 - 版本 - 许可证 - Issues - Discord - Ko-fi - PayPal -

- -

- 下载最新版本 + English

-
+![Electron 应用演示:声音克隆、声音设计、视频配音和模型管理](docs/media/electron/voicestudio.gif) -
- VoiceStudio — 从状态栏快速切换 TTS 引擎 -
+

新 Electron 桌面界面,使用此分支及内置演示声音录制。正式发布版本的界面可能有所不同。

-> **声音很私人,创作空间也应该真正属于你。** VoiceStudio 的核心流程运行在你的硬件上:克隆、设计、配音、听写,并以 646 种语言创作,不需要订阅,也没有用量计费。联网引擎和服务始终是清晰可见的可选项,而不是隐藏依赖。 +## 用 VoiceStudio 创作 -> [!WARNING] -> **活跃 Beta 阶段。** 各版本之间可能出现故障——如需最新修复,请从源码运行。非常欢迎 Bug 报告和 PR:[提交 Issue](https://github.com/debpalash/VoiceStudio/issues) 或 [加入 Discord](https://discord.gg/bzQavDfVV9)。 +- **声音克隆与设计**:上传参考录音,或用文字描述你想要的声音。 +- **视频配音**:转录、翻译、分配说话人,并编辑语音时间轴。 +- **语音听写**:通过悬浮录音组件录制、转录和复制文字。 +- **长篇创作**:制作多角色脚本、有声书和批量任务。 +- **模型管理**:选择语音合成与转录引擎、语言及计算设备。 - +本地工作流在你的硬件上运行。远程服务为可选功能;使用情况分析须经同意才会启用。 -## ⚡ 快速开始 + + + + + + +
Electron 声音克隆工作区与内置演示声音Electron 视频配音工作区
声音克隆视频配音
-
- 下载 macOS DMG - 下载 Windows MSI - 下载 Linux AppImage -
- 三个按钮都会打开最新发布页——在资源列表中下载对应你系统的安装包。
- macOS:首次启动需要一次性批准——右键点击 → 打开(macOS 15 上为 系统设置 → 隐私与安全性 → “仍要打开”)。无需终端。为什么? · Intel Mac:不支持本地后端(#889)——详情 -
+## 开始使用 -选择你的操作系统,按指南从头到尾操作: +从 [Releases](https://github.com/debpalash/VoiceStudio/releases/latest) 下载,然后阅读对应平台的安装指南: -- 🍎 **macOS** — [docs/install/macos.md](docs/install/macos.md) -- 🪟 **Windows** — [docs/install/windows.md](docs/install/windows.md) -- 🐧 **Linux** — [docs/install/linux.md](docs/install/linux.md) -- 🐳 **Docker** — [docs/install/docker.md](docs/install/docker.md) · [Docker Hub: `palashdeb/omnivoice-studio`](https://hub.docker.com/r/palashdeb/omnivoice-studio) +**[macOS](docs/install/macos.md) · [Windows](docs/install/windows.md) · [Linux](docs/install/linux.md) · [Docker](docs/install/docker.md)** + +打开声音克隆页面,选择已有声音或添加清晰的参考录音,输入文字并生成。按提示安装所需模型。硬件要求因引擎而异,详见[性能指南](docs/performance.md)。 + +**从源码运行 Electron 预览版:** ```bash -# Docker 快速运行 (CPU / 本地环回模式) -docker run -d -p 127.0.0.1:3900:3900 -v omnivoice-data:/app/omnivoice_data --name voicestudio palashdeb/omnivoice-studio:stable +git clone https://github.com/debpalash/VoiceStudio.git +cd VoiceStudio +bun install +cd electron +bun run dev ``` -**三步克隆出你的第一个声音:** +环境要求和后端配置见 [Electron 开发指南](electron/README.md)。项目仍在积极开发中,可通过 [GitHub Issues](https://github.com/debpalash/VoiceStudio/issues) 反馈问题。 -1. **安装并启动。** 首次启动会自动搭建 Python 运行环境并下载模型权重——启动画面会逐步显示进度(仅首次,需要几分钟;之后即开即用)。 -2. 从启动台打开**语音克隆**,拖入任意声音的 **3 秒音频**。 -3. **输入一句话,点击生成。** 音频在你的设备上生成并保存,支持 646 种语言(商业使用前请审阅所选模型与分词器的许可条款)。 +## 文档 -### 🎧 音频示例 - -在线试听 VoiceStudio 本地生成的实际音频样例: - -| 工作流 | 提示词 / 参考音频 | 生成音频 | -|---|---|---| -| **声音克隆** | [demo_voice.wav](backend/assets/samples/demo_voice.wav) | [demo_clone_output.wav](backend/assets/samples/demo_clone_output.wav) | -| **声音设计** (美语新闻主播) | *"清晰、权威的美国广播级音色"* | [demo_voice_design_us_news_anchor.wav](backend/assets/samples/voice_design/demo_voice_design_us_news_anchor.wav) | -| **声音设计** (英式有声书) | *"温暖生动的英式故事讲述音色"* | [demo_voice_design_audiobook_uk_narrator.wav](backend/assets/samples/voice_design/demo_voice_design_audiobook_uk_narrator.wav) | -| **视频配音** (多语种) | [source.src.wav](backend/assets/samples/demo/dubbing/source.src.wav) | [西班牙语](backend/assets/samples/demo/dubbing/dubbed_es.src.wav) · [法语](backend/assets/samples/demo/dubbing/dubbed_fr.src.wav) · [日语](backend/assets/samples/demo/dubbing/dubbed_ja.src.wav) · [中文](backend/assets/samples/demo/dubbing/dubbed_zh.src.wav) | - -觉得慢?[docs/performance.md](docs/performance.md) 讲清了生成时间到底花在哪里、有哪些调优开关,以及“它变慢了”的三个经典原因。各引擎/设备的实测数据见 [docs/benchmarks.md](docs/benchmarks.md)。 - -> 正在从 **[CorentinJ/Real-Time-Voice-Cloning](https://github.com/CorentinJ/Real-Time-Voice-Cloning)**(现已归档)迁移过来?我们有专门的迁移指南:[docs/migration/real-time-voice-cloning.md](docs/migration/real-time-voice-cloning.md)。 - -
-🧰 卡住了?自检、Token 与受限网络 - -
- -先运行内置自检——在应用中打开 **设置 → 关于 → “运行自检”**,或在源码检出目录中执行 -`uv run python backend/main.py --diagnose`(加 `--deep` 还会实际加载当前引擎进行测试)。然后查看 -[docs/install/troubleshooting.md](docs/install/troubleshooting.md) 中排名前 -10 的安装错误。运行时出错时,应用内的错误界面会直接深链到对应条目;**设置 → 关于 → -“保存诊断包”** 会把脱敏日志与自检报告打包,方便附在 Bug 报告里。 - -Hugging Face Token 的配置见 -[docs/setup/huggingface-token.md](docs/setup/huggingface-token.md)。说话人分离相关的模型访问门槛见 -[docs/features/diarization.md](docs/features/diarization.md)。下载速度、⚡ 快速下载(Xet)状态,以及受限网络 / 镜像选项见 -[docs/downloading-models.md](docs/downloading-models.md)。 - -
- ---- - - - -## ✨ 功能 - -八大主打功能——折叠区里还有十二项等你展开。 - - - - - - - - - - - - - - -
-

🎙️ 语音克隆

-

3 秒音频 → 复刻任何声音。
646 种语言,零样本。

-
-

🎨 声音设计

-

性别、年龄、口音、音高、语速、
情感、方言——随心调节

-
-

🎬 视频配音

-

YouTube 链接或文件 → 转录 →
翻译 → 重新配音 → MP4

-
-

📖 有声书编辑器

-

导入文本、EPUB 或 PDF。自动分章、
响度归一、元数据。导出 .m4b

-
-

🎭 故事模式

-

多声音编辑器。逐行分配声音、
预览、导出完整配音阵容

-
-

⌨️ 听写工具

-

任何应用中按 ++Space
转录、自动粘贴、随即消失。

-
-

🔐 本地优先

-

核心创作流程
留在你的设备上

-
-

🤖 MCP 服务器

-

Claude、Cursor 或
任何 MCP 客户端使用 VoiceStudio。

-
- -
-……还有 12 项——人声分离、说话人分离、批量处理、水印、诊断等等 - -
- -- 🔊 **人声分离** — 基于 Demucs:把语音从音乐中分离出来,同时保留背景音床。 -- 👥 **说话人分离** — Pyannote + WhisperX 自动识别谁说了什么。 -- 📦 **批量队列** — 拖入 50 个视频就可以走开;每个任务都有独立进度条。 -- 🛡️ **AI 水印** — AudioSeal(Meta):不可见,且能在压缩后留存。 -- 🔬 **诊断** — 自检套件、错误日志、脱敏诊断包。 -- ⚡ **GPU 自动检测** — CUDA · MPS · ROCm(Linux,需手动开启)· CPU;显存 ≤8 GB 时自动卸载。 -- 🧭 **引擎路由** — 逐引擎 GPU 预检;绝不静默回退到 CPU。 -- 🧩 **可扩展** — 继承 `TTSBackend`,约 50 行代码即可接入任意引擎。 -- 🎒 **便携声音角色** — 将声音导出为 `.ovsvoice` 包:身份 + 水印。 -- ♾️ **无限长 TTS** — 按句分块生成,没有长度上限,可经 WebSocket 流式输出。 -- 🌐 **远程后端** — 让 UI 指向远程服务器;对 Tailscale 友好,支持 Bearer 认证。 -- 🧠 **听写 + LLM** — 用本地 LLM 润色转录文本,可选回声消除。 - -
- ---- - - - -## 💡 为什么选择 VoiceStudio? - -云端语音工具很方便,但工作流会依赖账号、用量计费和他人的基础设施。VoiceStudio 在你的硬件上提供完整工作室;只有你主动选择时,才会使用联网集成。 - -| | **ElevenLabs** | **VoiceStudio** | -|---|---|---| -| **价格** | 订阅与用量限制 | 免费且开源(AGPL-3.0)· 专有用途可选 [商业许可证](#license) | -| **语音克隆** | ✅ 3 秒音频 | ✅ 3 秒音频,零样本 | -| **声音设计** | ✅ 性别、年龄 | ✅ 性别、年龄、口音、音高、风格、方言 | -| **有声书 / 故事** | ❌ | ✅ 完整有声书编辑器 + 多声音故事(EPUB/PDF 导入,.m4b 导出) | -| **语言** | 取决于套餐和模型 | **646** | -| **视频配音** | ✅ 仅云端 | ✅ 完全本地 | -| **数据隐私** | 音频在远端处理 | 核心流程在本地运行;联网服务必须主动选择 | -| **API 密钥** | 需要账号 | 本地流程不需要 | -| **GPU 支持** | 不适用(云端) | CUDA · Apple Silicon · ROCm(Linux)· CPU | -| **桌面应用** | ❌ | ✅ macOS · Windows · Linux | -| **TTS 引擎** | 1 | **16** — [完整矩阵](#tts-engines) | -| **ASR 引擎** | 1 | **11** — [完整阵容](#asr-engines) | -| **MCP 服务器** | ❌ | ✅ 可从 Claude、Cursor 及任何 MCP 客户端使用 | -| **自检** | ❌ | ✅ 诊断套件、错误日志、脱敏调试包 | -| **可定制** | ❌ 闭源 | ✅ 随你 Fork、扩展、发布 | - -专业级语音 AI,去掉订阅,也去掉云端。 - -
-
- 心动了?来和我们一起构建吧。
- 加入 Discord -

-
- ---- - -## 🖥️ 系统要求 - -| | **最低配置** | **推荐配置** | -|---|---|---| -| **操作系统** | Windows 10、macOS 12+(Apple Silicon)、Ubuntu 24.04+(glibc 2.39+) | 任意现代 64 位操作系统 | -| **内存** | 8 GB | 16 GB+ | -| **显存(GPU)** | 4 GB(自动将 TTS 卸载到 CPU) | 8 GB+(NVIDIA RTX 3060+) | -| **硬盘** | 10 GB 可用空间(模型 + 缓存) | 20 GB+ SSD | -| **Python** | 3.10+(由 `uv` 管理) | 3.11–3.12 | -| **GPU** | 可选——CPU 也能跑 | NVIDIA CUDA · Apple Silicon MPS · AMD ROCm(仅 Linux) | - -> [!TIP] -> 对于显存 **≤8 GB** 的 GPU,VoiceStudio 会在转录期间自动将 TTS 卸载到 CPU——无需配置。不需要专用 GPU;整条流水线都可以在 CPU 上运行(只是慢一些)。 - -> [!NOTE] -> **AMD GPU:** ROCm 加速**仅限 Linux 且需手动开启**——在首次运行的设置界面选择 **“AMD GPU (ROCm)”**,或设置 `OMNIVOICE_TORCH_VARIANT=rocm`([docs/install/linux.md](docs/install/linux.md#amd-gpu-rocm))。在 **Docker/Podman** 中请改用专门的 ROCm 镜像:`ghcr.io/debpalash/omnivoice-studio:rocm`([docs/install/docker.md](docs/install/docker.md#pull-and-run-amd-gpu--rocm))。**在 Windows 上,AMD GPU(含 Ryzen AI 核显)只能以 CPU 运行**:PyTorch 没有 Windows 版 ROCm 轮子,因此 Windows 上的 GPU 加速仅限 NVIDIA/CUDA([docs/install/windows.md](docs/install/windows.md#gpu-support))。 - -> [!IMPORTANT] -> **macOS Intel(x86_64)不支持本地后端:** 应用 UI 可以安装,但 Python 后端无法运行,因为 PyTorch 已不再发布 Intel Mac 轮子([#889](https://github.com/debpalash/VoiceStudio/issues/889))。Intel Mac 用户仍可让 UI 指向另一台机器上的远程后端——参见 [docs/install/macos.md](docs/install/macos.md)。 - - - -### 💡 按硬件推荐引擎配置 - -| 硬件配置 | 推荐 TTS 引擎 | 推荐 ASR 语音识别 | 优势 | -|---|---|---|---| -| **Apple Silicon (M1–M4)** | [MLX-Audio](docs/engines/mlx-audio.md) · [OmniVoice](docs/engines/omnivoice.md) (MPS) | [MLX Whisper](docs/engines/mlx-whisper.md) · [Parakeet MLX](docs/engines/parakeet-mlx.md) | 原生统一内存,macOS 上延迟最低、性能最强 | -| **NVIDIA 显卡 (8 GB+ 显存)** | [OmniVoice](docs/engines/omnivoice.md) · [CosyVoice 3](docs/engines/cosyvoice.md) | [WhisperX](docs/engines/whisperx.md) | 极致零样本克隆品质、字级时间戳对齐与说话人分离 | -| **低显存 / 仅 CPU 设备** | [PocketTTS](docs/engines/pockettts.md) · [Sherpa-ONNX](docs/engines/sherpa-onnx.md) · [KittenTTS](docs/engines/kittentts.md) | [Moonshine](docs/engines/moonshine.md) · [Faster-Whisper](docs/engines/faster-whisper.md) (`int8`) | 超低内存占用,针对 CPU 指令集深度优化 | - - - -### 🗣️ TTS 引擎 - -**16 个引擎,一个选择器。** VoiceStudio(默认,支持 600+ 语言)始终可用;另有七个引擎可选装并自动检测(CosyVoice 3、GPT-SoVITS、VoxCPM2、MOSS-TTS-Nano、KittenTTS、MLX-Audio、Sherpa-ONNX),外加八个按需延迟安装的引擎(IndexTTS 2.5、OmniVoice GGUF、OmniVoice 子进程版、PocketTTS、Supertonic 3、MOSS-TTS-v1.5、dots.tts、Confucius4-TTS)。在 **设置 → TTS 引擎** 中切换;所选引擎将应用于所有语音合成场景。**每个引擎都有独立指南:[docs/engines](docs/engines/README.md)(英文)。** - -
-📊 完整矩阵——16 个引擎 × 平台 × 克隆/指令 × 许可证 - -
- -| 引擎 | 语言 | 克隆 | 指令 | Linux | macOS ARM | Windows | 许可证 | -|--------|:---------:|:-----:|:--------:|:-----:|:---------:|:-------:|:-------:| -| **VoiceStudio**(默认,由 k2-fsa/OmniVoice 驱动) | 600+ | ✅ | ✅ | ✅ CUDA/CPU | ✅ MPS | ✅ CUDA/CPU | 内置 | -| **CosyVoice 3** | 9 + 18 种方言 | ✅ | ✅ | ✅ CUDA/CPU | ✅ MPS | ✅ CUDA/CPU | Apache-2.0 | -| **GPT-SoVITS** | 5 | ✅ | — | ✅ CUDA/CPU | — | ✅ CUDA/CPU | MIT | -| **VoxCPM2** | 30 | ✅ | ✅ | ✅ CUDA/CPU | ✅ MPS | ✅ CUDA/CPU | Apache-2.0 | -| **MOSS-TTS-Nano** | 20 | ✅ | — | ✅ CUDA/CPU | ✅ CPU | ✅ CUDA/CPU | Apache-2.0 | -| **KittenTTS** | 英语 | — | — | ✅ CPU | ✅ CPU | ✅ CPU | MIT | -| **MLX-Audio**(Kokoro、Qwen3-TTS、CSM、Dia 等) | 多语言 | 因模型而异 | 因模型而异 | ❌ | ✅ 原生 | ❌ | 因模型而异 | -| **Sherpa-ONNX** | 20+ | — | — | ✅ CUDA/CPU | ✅ CPU | ✅ CUDA/CPU | Apache-2.0 | -| **IndexTTS 2.5** ⚡ | 中文 · 英语 · 日语 · 西班牙语 · 阿拉伯语 | ✅ | — | ✅ CUDA | — | ✅ CUDA | Bilibili 模型许可¹ | -| **OmniVoice GGUF** ⚡ | 600+ | ✅ | ✅ | ✅ CPU | ✅ CPU | ✅ CPU | 内置 | -| **Supertonic 3** ⚡ | 31 | — | — | ✅ CPU | ✅ CPU | ✅ CPU | OpenRAIL-M | -| **MOSS-TTS-v1.5** ⚡(8B) | 31 | ✅ | — | ✅ CUDA/CPU | ✅ CPU | ✅ CUDA/CPU | Apache-2.0 | -| **dots.tts** ⚡(2B) | 24 | ✅ | — | ✅ CUDA/CPU | ✅ CPU | ❌ | Apache-2.0 | -| **Confucius4-TTS** ⚡ | 14 | ✅ | — | ✅ CUDA/CPU | ✅ CPU | ✅ CUDA/CPU | Apache-2.0 | - -¹ 若月活跃用户超过 1 亿,或年收入超过人民币 10 亿元,使用 IndexTTS 2.5 -前必须另行取得 Bilibili 的书面许可。启用可选边车前,请审阅其 -[模型许可](https://huggingface.co/IndexTeam/IndexTTS-2.5/blob/main/LICENSE)。 - -> **CUDA** = GPU 加速 · **MPS** = Apple Silicon Metal · **CPU** = 随处可运行,大模型较慢 · KittenTTS 和 MOSS-TTS-Nano 可在 CPU 上实时运行 · MLX-Audio 仅限 Apple Silicon · ⚡ = 延迟注册(首次使用时安装) -> -> **克隆**能力的意义不止于单段生成:视频配音(以及任何固定了声音的批量任务)需要参考音频克隆来保持说话人身份,因此把不支持克隆的引擎(KittenTTS、Sherpa-ONNX、Supertonic 3)设为当前引擎时,这些任务会在开始前就给出可操作的失败提示,而不是静默回退到 VoiceStudio。 -> -> **MOSS-TTS-v1.5**(8B,约 16 GB)、**dots.tts**(2B,约 9 GB)和 **Confucius4-TTS** 是重量级可选引擎,从本地克隆在各自独立的 venv 中运行。三者均不支持 Apple Silicon MPS(在 Mac 上以 CPU 运行);dots.tts 没有 Windows 路径;Confucius4 建议使用 CUDA(CPU 可用,约为实时时长的 17 倍)。详情:[MOSS-TTS-v1.5](docs/engines/moss-tts-v15.md) · [dots.tts](docs/engines/dots-tts.md) · [Confucius4-TTS](docs/engines/confucius4-tts.md)。 - -
- - - -### 🎧 ASR 引擎 - -**11 个引擎**——它们驱动听写、视频配音和字幕。**WhisperX** 是跨平台的默认引擎(约 100 种语言,词级时间对齐);其余引擎均为可选装并自动检测。在 **设置 → 引擎** 中切换。十个完全在本地设备上运行;第十一个(OpenAI 兼容)是可选的远程客户端,可用于 Qwen3-ASR 或任何兼容的服务器。 - -
-📊 完整阵容——11 个引擎、各自的强项与计算类型说明 - -
- -| 引擎 | `OMNIVOICE_ASR_BACKEND` | 语言 | 最适合 | -|--------|-------------------------|:---------:|----------| -| **WhisperX**(默认) | `whisperx` | ~100 | 配音与字幕——通过 wav2vec2 强制对齐实现词级时间对齐 | -| **Faster-Whisper** | `faster-whisper` | ~100 | Linux / macOS / Windows 上的快速转录(CTranslate2) | -| **Faster-Whisper(隔离)** | `faster-whisper-isolated` | ~100 | 与 Faster-Whisper 相同,但在子进程中崩溃隔离——ASR 崩溃不会拖垮整个应用 | -| **MLX Whisper** | `mlx-whisper` | ~100 | Apple Silicon 原生速度(Apple MLX / Metal) | -| **PyTorch Whisper** | `pytorch-whisper` | ~100 | 经 🤗 Transformers 的 CUDA / CPU 兜底方案(无需 cuDNN 8) | -| **Parakeet TDT** | `nemo-parakeet` | 英语 + 25 种欧洲语言 | 即使在 CPU 上也能以约 10 倍实时速度达到 SOTA 精度,自动语言检测(NVIDIA NeMo,CUDA/CPU) | -| **Moonshine** | `moonshine` | 英语 | 边缘设备 / 低延迟,ONNX | -| **FunASR** | `funasr` | 50+ | 多语言一体化——内置 VAD + 行内说话人分离(SenseVoice) | -| **sherpa-onnx**(实时听写) | `sherpa-onnx-asr` | 25 种欧洲语言 + 90+ | 实时、快于实时的听写——小体积流式/离线 ONNX 模型(Parakeet TDT v3/v2、流式 Zipformer 与 Paraformer、Whisper Tiny),CPU 运行,macOS / Windows / Linux 表现完全一致。在 **设置 → 语音** 中按模型选择。 | -| **OpenAI 兼容** ⚠️ 远程 | `openai-compat-asr` | 取决于服务器 | 当下通往 **Qwen3-ASR** 的路径(自托管服务器,无需等 transformers 支持)、任何 OpenAI 兼容的转录端点,或 OpenAI 官方 API——无需安装,在 **设置 → 引擎**(ASR 标签页)中配置并测试连接。音频会离开你的设备,发送到你指定的任何服务器;参见 [docs/engines/openai-compatible-asr.md](docs/engines/openai-compatible-asr.md)。 | - -> Whisper 系列引擎覆盖约 100 种语言;**FunASR / SenseVoice** 额外提供一条多语言一体化路径,内置语音活动检测与行内说话人分离。**sherpa-onnx** 驱动实时听写的模型选择器——你边说,文字边出现。除可选的 OpenAI 兼容远程客户端外,所有引擎都在本地设备上运行——无需 API 密钥,无需云端。 - -> **GPU 不支持高效 float16?** 在较老的 NVIDIA GPU(Maxwell/Pascal、GTX 16xx)上,或在 CTranslate2/cuDNN 版本不匹配之后,CTranslate2 系 ASR 引擎(WhisperX、Faster-Whisper)无法运行 `float16`,VoiceStudio 会自动改用 `int8` 重试——无需配置。如果转录仍然失败,可用 `ASR_COMPUTE_TYPE` 环境变量固定计算类型(逃生舱口):`ASR_COMPUTE_TYPE=int8`(CPU 用 `float32`)。将其设为 `int8` 并重启后端。 - -
- ---- - -## 🏗️ 架构 - -``` -┌─────────────────────────────────────────────────────────────┐ -│ Frontend (React) │ -│ DubTab · VoiceConsole · Stories · Audiobook · Gallery │ -│ Dictation · BatchQueue · Diagnostics · MCP Client │ -├─────────────────────────────────────────────────────────────┤ -│ Backend (FastAPI) │ -│ 100+ API endpoints · SSE+WSS streaming · SQLite │ -├──────────┬──────────┬──────────┬──────────┬────────────────┤ -│ WhisperX │ Demucs │VoiceStudio │ Pyannote │ Engine Routing │ -│ (+7 ASR │ Source │ (+10 │ Diariz- │ ↳ GPU preflight │ -│ engines) │ Sep. │ TTS) │ ation │ ↳ No silent CPU │ -└──────────┴──────────┴──────────┴──────────┴────────────────┘ - CUDA / MPS / ROCm / CPU (auto-detected + routed) -``` - - - -## 🔌 OpenAI 兼容 API - -已经有会说 OpenAI 音频 API 的脚本、智能体或工具?把它指向 `http://localhost:3900/v1` 即可——不需要密钥,也不用改代码。后端为音频端点内置了即插即用的兼容接口,直接接到你当前启用的 TTS/ASR 引擎(没错,`voice` 参数接受你克隆的声音配置 ID)。 - -| 端点 | 作用 | +| 需求 | 链接 | |---|---| -| `POST /v1/audio/speech` | TTS——输入文本;输出 `mp3` / `wav` / `flac` / `opus` / `pcm`。`tts-1` / `tts-1-hd` 映射到你当前启用的引擎;也接受 OpenAI 的声音名称(`alloy` 等)。 | -| `POST /v1/audio/transcriptions` | STT——输入音频文件;输出 `json`、`text`、`verbose_json`、`srt` 或 `vtt`。`whisper-1` 映射到你当前启用的 ASR 引擎。 | -| `GET /v1/audio/voices` | VoiceStudio 扩展——列出所有声音配置和引擎,客户端可据此发现你的克隆声音。 | +| 安装帮助 | [故障排查](docs/install/troubleshooting.md) · [模型下载](docs/downloading-models.md) | +| 模型与音质 | [引擎指南](docs/engines/README.md) · [基准测试](docs/benchmarks.md) | +| 集成 | [本地 API](docs/speech-platform.md) · [MCP](docs/mcp.md) · [示例](examples/README.md) | +| 参与开发 | [贡献指南](.github/CONTRIBUTING.md) · [Electron](electron/README.md) · [更新日志](CHANGELOG.md) | -```sh -curl http://localhost:3900/v1/audio/speech \ - -H "Content-Type: application/json" \ - -d '{"model": "tts-1", "voice": "alloy", "input": "Generated on my own hardware.", "response_format": "wav"}' \ - --output speech.wav -``` +安装智能体技能:`npx skills add debpalash/VoiceStudio` -```python -from openai import OpenAI -client = OpenAI(base_url="http://localhost:3900/v1", api_key="none") # any string works — nothing checks it +## 支持 VoiceStudio -result = client.audio.transcriptions.create(model="whisper-1", file=open("clip.wav", "rb")) -print(result.text) -``` +[Ko-fi](https://ko-fi.com/debpalash) · [PayPal](https://paypal.me/palashCoder) · [赞助项目](SPONSORS.md) · [商务合作](mailto:partner@voicestudio.sh) -想要完整的接口(100+ 端点)?完整的 REST API 参考已内嵌在应用中——**设置 → OpenAPI 参考**(由 Scalar 驱动),或点击页脚的 `{}` 按钮。 +## 许可与负责任使用 -### 📓 在 Google Colab 上运行 - -[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/debpalash/VoiceStudio/blob/main/notebooks/OmniVoice_Studio_Colab.ipynb) - -没有本地 GPU?官方笔记本([notebooks/OmniVoice_Studio_Colab.ipynb](notebooks/OmniVoice_Studio_Colab.ipynb))可在免费的 Colab T4 上启动完整应用(包含 Web 界面):在笔记本内直接构建前端,用 uv 安装后端(复用 Colab 预装的 CUDA PyTorch),并通过 Colab 内置端口代理打开界面。无需第三方隧道,也无需任何 API 密钥。随后还有一套覆盖全部主要功能的 API 导览,全部可在笔记本内直接播放:多语言 TTS、声音克隆与声音设计、已保存的声音档案、语音转写、AI 水印检测、OpenAI 兼容 API、多角色故事、带章节的 m4b 有声书,以及一个附带人声分离音轨的迷你视频配音。 - -### 🤝 智能体技能(Agent Skills) - -用一条命令教会你的 AI 智能体(Claude Code、Cursor、Codex 等)使用 VoiceStudio: - -```sh -npx skills add debpalash/omnivoice-studio -``` - -内含两个 [skills](https://skills.sh):**`omnivoice`**——让任何智能体通过你的本地安装进行语音合成与转录(包括你克隆的声音),免费且离线;以及 **`oss-maintainer`**——本项目所遵循的维护者方法论,适合任何用智能体运营自己开源项目的人。 - -### 🔌 模型上下文协议(MCP 服务器) - -VoiceStudio 在 `http://localhost:3900/mcp` 挂载了 MCP 服务,可供 Claude Desktop、Cursor 与自主智能体调用: - -```json -{ - "mcpServers": { - "voicestudio": { - "url": "http://localhost:3900/mcp" - } - } -} -``` - -对于需要 stdio 管道传输的客户端,请使用内置的本地桥接脚本(`docs/mcp.json`): - -```json -{ - "mcpServers": { - "voicestudio": { - "command": "python", - "args": ["-m", "backend.mcp_shim"], - "cwd": "/path/to/VoiceStudio" - } - } -} -``` - -支持 `generate_speech`、`clone_voice`、`transcribe` 等工具与流式文件输出模式,详见 [docs/mcp.md](docs/mcp.md)。 - ---- - -## 🗺️ 路线图 - -### 🔜 即将推出 - -- 🎬 **唇形同步 v2** — 使用 wav2lip 进行视觉语音时间对齐 -- 🌐 **在线演示** — 无需安装即可体验 VoiceStudio -- 🔌 **插件市场** — 社区贡献的 TTS 引擎与特效 -- 🎵 **实时变声器** — 通话中的麦克风实时变声 - -
-✅ 已经发布的一切——按类别列出的“成绩单” - -
- -| 分类 | 功能 | -|----------|----------| -| **长内容** | 有声书编辑器(文本/EPUB/PDF → 分章 .m4b)、Stories 多声音编辑器、两遍响度归一母带处理、渲染中断后的崩溃续渲、发音控制 + SSML-lite 韵律 | -| **配音** | 完整流水线(转录→翻译→合成→封装)、场景感知分割、唇形同步评分、流式 TTS、逐说话人声音分配、Smart Fit 时长匹配 + 二次 QC、独立的配音主页 | -| **声音** | 零样本克隆、声音设计、A/B 对比、声音预览控件、支持收藏/标签的声音库、便携声音角色包(`.ovsvoice`)、声音控制台工作区 | -| **音频** | Demucs 人声分离、逐段增益、选择性音轨导出、分轨/SRT/VTT/MP3 导出、按句分块实现的无限长 TTS | -| **多语言** | 多语言批量选择器、顺序 GPU 执行的批量配音队列 | -| **说话人分离** | Pyannote 机器学习分离、自动说话人克隆提取、逐说话人声音分配 | -| **ASR** | 9 个引擎(WhisperX、Faster-Whisper、隔离版 Faster-Whisper、MLX Whisper、PyTorch Whisper、Parakeet TDT、Moonshine、FunASR/SenseVoice、sherpa-onnx 实时听写)、崩溃隔离的子进程后端 | -| **TTS** | 14 个引擎(VoiceStudio、CosyVoice 3、GPT-SoVITS、VoxCPM2、MOSS-TTS-Nano、KittenTTS、MLX-Audio、Sherpa-ONNX,+ 延迟安装:IndexTTS 2.5、OmniVoice GGUF、Supertonic 3、MOSS-TTS-v1.5、dots.tts、Confucius4-TTS)、带 GPU 预检的引擎路由 | -| **基础设施** | Docker 部署、CUDA/MPS/ROCm 自动检测、cuDNN 8 兼容、显存感知模型卸载、引擎路由(绝不静默回退 CPU)、诊断套件与错误日志、受限网络镜像支持 | -| **AI 溯源** | AudioSeal 不可见水印(类似 SynthID)、视频徽标叠加、水印检测 API | -| **用户体验** | 撤销/重做、键盘快捷键、拖放、会话持久化、首次启动按屏幕推荐界面缩放,以及原生 WebKitGTK 缩放 | -| **实时事件** | WebSocket 事件总线——数据变更时即时刷新侧边栏、指数退避重连 | -| **状态管理** | Zustand 状态迁移——`uiSlice`、`pillSlice`、`dubSlice`、`generateSlice`、`prefsSlice`、`glossarySlice` | -| **桌面** | 跨平台 Tauri 安装程序(macOS DMG——Apple Silicon;Intel 不支持本地后端,#889——Windows MSI、Linux deb/AppImage)、自动更新基础设施、单实例约束、关闭最小化到托盘、macOS Gatekeeper 修复 | -| **听写** | 全局系统级热键(`⌘+⇧+Space`)、无边框浮动控件、WebSocket 流式 ASR、自动粘贴、可自定义热键、本地 LLM 转录润色 | -| **批量流水线** | 完整批量 TTS:提取 → 转录 → 翻译 → 生成 → 混音 → 导出,带实时进度追踪 | -| **MCP 服务器** | 让 VoiceStudio 成为 Claude、Cursor 及任何 MCP 客户端的本地 TTS/STT 提供方 | -| **远程后端** | 让桌面 UI 指向远程后端 URL,支持 Bearer 认证(附 Tailscale 文档) | -| **可靠性** | 启动开屏的卡死看门狗、逐引擎 GPU 兼容矩阵、引擎二进制不可执行时的可操作报错、setuptools 自动修复 | - -
- ---- - - - -## 💜 赞助 / 捐赠 - -VoiceStudio 由一位开发者使用 Claude Code 和 AI 智能体独立打造——而智能体账单是实打实的(过去三个月花了数千美元)。如果 VoiceStudio 为你创造了价值,帮忙分担一小部分账单,就能让开发保持全职推进。 - -
- -**本月智能体账单基金** - -已筹 $10 / $200 - -

- -Ko-fi -   -PayPal - -
-每一美元都直接用于支付智能体账单——让 VoiceStudio 的开发持续不断。 - -

- -来自 VoiceStudio 作者的更多应用——同样的本地优先理念: -Opal 💠(播放一切——AI 时代的媒体播放器)· -memxt 🧠(Claude Code 与编码智能体的本地记忆)。 -给它们点个 ⭐ 也是一种支持 → 详见下文 - -
- - - -### 🌟 赞助商 - -VoiceStudio **免费**且采用 **AGPL-3.0** 许可——没有付费版,没有 SaaS 收入。赞助商让开发得以持续,作为回报,可以在这里、在应用内(顶级档位还包括项目官网)获得一个徽标位。这是一份感谢,绝不是付费墙。**[查看档位并成为赞助商 →](SPONSORS.md)** - -
- - - -**这里可以是你的徽标** — [成为赞助商](SPONSORS.md) - - - -
- -💡 GitHub 也会在本仓库顶部显示一个 **Sponsor** 按钮,经由 .github/FUNDING.yml 指向相同的链接。 - ---- - -## 💬 社区 - -
- 加入 Discord -
- 设置类问题我们几小时内就会回复,而不是几天。 -
- -
-里面都在聊什么 - -
- -| 频道 | 那里发生什么 | -|---------|--------------------| -| `#announcements` | 发布消息与重大时刻——新版本最先在这里公布 | -| `#releases` + `#changelog` | 每一个构建,以及里面究竟有什么 | -| `#issues` | 以论坛帖子形式提交的 Bug 报告——直接分诊进 GitHub Issues | -| `#ideas` | 功能请求,供讨论与投票 | -| `#discuss-ideas` | 动手之前的设计讨论 | -| `#general` | 安装帮助、GPU 疑难排查,以及晒你的配音成果 | - -
- ---- - - - -## 🤝 参与贡献 - -非常欢迎——Bug 修复、新的 TTS 引擎适配器、UI 改进、文档、翻译。统统欢迎。 - -- 📖 阅读 **[贡献指南](.github/CONTRIBUTING.md)** 了解环境搭建、代码风格和 PR 工作流 -- 🐛 浏览 [good first issues](https://github.com/debpalash/VoiceStudio/labels/good%20first%20issue) -- 💬 加入我们的 [Discord](https://discord.gg/bzQavDfVV9) 讨论想法或寻求帮助 - ---- - -## ❓ 常见问题 - -
-真的能和 ElevenLabs 一样好吗? -
-诚实的回答:取决于你要做什么。 - -VoiceStudio 真正有竞争力的地方:从干净的参考音频进行语音克隆(最先进的开源扩散 TTS)、语言覆盖(646 种语言对他们的 32 种),以及所有结构性优势——没有按字符计费、没有用量上限、音频不离开你的设备、完整的流水线可定制性(14 个 TTS 引擎、10 个 ASR 引擎、翻译方案随你选)。 - -ElevenLabs 仍然领先的地方:开箱即用的稳定性与打磨程度,尤其是英语 TTS。他们的单一模型经过深度调优;我们的质量取决于你选择的引擎、你的硬件,以及(对克隆而言)参考音频——干燥、近麦的音频比嘈杂或有回声的音频克隆效果好得多。 - -具体到配音:配音是一条链——转录 → 翻译 → 克隆 → 合成——在你的素材上,它只取决于最薄弱的一环。如果部分输出语无伦次,先检查片段表里的原文:当转录本身就错了,换一个 ASR 引擎或使用更干净的源音频——修复点通常在这里,而不是声音。 - -拿你的真实素材试试——免费,下载一次即可。许多用户直接用它替换了 ElevenLabs;也有人两个都留着。这两种结果我们都乐见。 -
- -
-能在 Apple Silicon(M1/M2/M3/M4)上运行吗? -
-可以。MPS 加速会被自动检测。在 Apple 硬件上,MLX 优化的 Whisper 模型可提供更快的转录速度。不支持 Intel Mac:应用 UI 可以安装,但本地 Python 后端无法运行,因为 PyTorch 已不再发布 Intel Mac 轮子(#889)——Intel Mac 只能配合远程后端使用。 -
- -
-需要多少显存? -
-最低 4 GB。 显存 ≤8 GB 时,TTS 模型会在转录期间自动卸载到 CPU。8 GB 以上时,所有组件同时在 GPU 上运行。完全没有 GPU?CPU 模式也能用——只是慢一些(TTS 约慢 3 倍)。 -
- -
-可以用于商业用途吗? -
-可以——商业使用免费,基于 AGPL-3.0:运行它、出售用它生成的音频、为客户的视频配音、在团队中部署。只有一项义务:如果你修改了 VoiceStudio 并通过网络向他人提供该修改版本,你必须依据相同条款分享修改后的源代码。想把它嵌入闭源产品?可获取商业许可证——参见许可证。 -
- -
-支持哪些语言? -
-通过 VoiceStudio 模型的 TTS 支持 646 种语言。转录(WhisperX)支持 99 种语言。翻译覆盖范围取决于目标语言对。 -
- -
-可以添加自己的 TTS 引擎吗? -
-可以。在 backend/services/tts_backend.py 中继承 TTSBackend,并将其添加到 _REGISTRY 字典中——约 50 行代码。十四个内置引擎均以此方式实现;参见 TTS 引擎。 -
- -
-VoiceStudio 会收集我的任何数据吗? -
-除非你明确同意,否则不会。首次运行时应用会询问你——一个页面、两个同等分量的按钮,没有预先勾选。在你回答“是”之前,VoiceStudio 什么都不发送:没有分析、没有遥测、没有账号、没有“回传”。跳过提问就等于“否”。无论如何,你的文本、音频、声音和项目永远不会离开你的设备。 - -如果你选择同意(也可随时在 设置 → 隐私 → “帮助改进 VoiceStudio” 中开关),发送的只是匿名、不含内容的使用统计:生成信息(引擎、语言、生成耗时、字符数量、错误类型),以及应用生命周期——一次安装信号、版本更新(版本号之间)、崩溃(错误类别和分桶后的运行时长,绝不含日志)、错误类型(有上限、去重),以及卸载时的一次告别信号。绝不包含你的文本、音频、文件名或任何可识别信息——这由代码中的属性白名单强制保证(backend/core/analytics.py),而不只是一句承诺。源码构建根本没有分析数据的接收端,因此根本不会询问。你自己的统计数字在 设置 → 用量 中查看,本地计算,不发送到任何地方。 -
- -
-如何卸载它 / 删除它的所有数据? -
-VoiceStudio 完全本地运行——卸载就是删除应用及其写入的文件夹(模型缓存、Python 环境、你的声音/项目、配置)。运行 scripts/uninstall.sh(macOS/Linux)或 scripts\uninstall.ps1(Windows)——它会先以干跑方式列出每个文件夹及其大小,加 --yes 才会真正删除。完整的各平台路径列表和应用移除步骤见 docs/install/uninstall.md。 -
- -## 🛡️ 负责任使用与安全 - -VoiceStudio 在个人硬件上提供零样本语音克隆与语音创作能力。我们提倡负责任的技术使用: -- **明确授权:** 严禁在未经说话人本人知情并明确授权的情况下克隆其声音。 -- **AI 溯源:** VoiceStudio 默认集成 [AudioSeal](https://github.com/facebookresearch/audioseal) 不可见神经音频水印,在完全不影响听感音质的前提下精准标记合成语音。 -- **本地隐私:** 默认本地工作流下,所有音频、声音档案、项目与转录文本始终保存在你的本地设备上;仅当你主动配置远程工作节点或第三方 ASR 端点时,相应数据才会传输到对应服务。 - ---- - - - -## 📜 许可证 - -VoiceStudio 是基于 [**GNU Affero 通用公共许可证 v3.0(AGPL-3.0)**](https://www.gnu.org/licenses/agpl-3.0.html) 的自由开源软件。 - -**可免费用于任何用途——包括商业和企业内部用途。** 运行它、出售用它生成的音频、为自己或客户的视频配音、在团队中推广——全部免费,无需许可证。作为一份**网络著佐权(copyleft)**许可证,AGPL 增加了一项义务:如果你**修改**了 VoiceStudio 并通过网络向他人提供该修改版本,你必须依据相同的 AGPL-3.0 条款向他们提供该修改版本的完整对应源代码。 - -希望将 VoiceStudio 嵌入**闭源或专有**产品或服务、又不受 AGPL-3.0 著佐权义务约束的组织,可获取**商业许可证**。**定价方案即将推出。** 咨询:**VoiceStudio@palash.dev**。 - -捆绑的 `omnivoice/` TTS 模型(作者 Han Zhu)在上游仍为 Apache-2.0 许可。完整且具约束力的条款请参见 [`LICENSE`](LICENSE)。 - ---- - -## 🙏 致谢 - -VoiceStudio 站在这些杰出开源工作的肩膀上: - -| 项目 | 作用 | -|---------|------| -| [**VoiceStudio (k2-fsa)**](https://github.com/k2-fsa/OmniVoice) | 零样本扩散 TTS 引擎——核心语音合成模型 | -| [**WhisperX**](https://github.com/m-bain/whisperX) | 词级别语音识别与时间对齐 | -| [**Demucs (Meta)**](https://github.com/facebookresearch/demucs) | 音乐源分离,用于人声分离 | -| [**Pyannote**](https://github.com/pyannote/pyannote-audio) | 说话人分离——谁说了什么 | -| [**CTranslate2**](https://github.com/OpenNMT/CTranslate2) | CPU 和 GPU 上的优化 Transformer 推理 | -| [**AudioSeal (Meta)**](https://github.com/facebookresearch/audioseal) | 用于 AI 溯源的不可见神经音频水印 | -| [**Tauri**](https://tauri.app) | 原生桌面应用框架 | -| [**Supertone / Supertonic 3**](https://huggingface.co/Supertone/supertonic-3) | ONNX TTS 引擎——31 种语言,CPU 高效 | -| [**Sherpa-ONNX**](https://github.com/k2-fsa/sherpa-onnx) | 支持 WASM 的通用 TTS/ASR 运行时 | -| [**GPT-SoVITS**](https://github.com/RVC-Boss/GPT-SoVITS) | 零样本 TTS 引擎——5 种语言,RTF 0.014 | - ---- - - - -## 🧰 来自同一作者的更多本地开源项目 - -喜欢这种本地优先的理念?它是一脉相承的——同一位作者,同一条准则:**你的数据只留在你的设备上。** 全部项目见 [palash.dev](https://palash.dev)。 - - - - - - -
-
- Opal 徽标 -

Opal 💠

-

播放一切。AI 时代的媒体播放器。

-

视频、动漫、漫画、种子、Jellyfin 和 Plex——一个播放器全部搞定,并内置本地 AI 记忆与上下文。使用 Zig 编写,支持 macOS 和 Windows。

-

- Opal Star 数 - Opal 官网 -

-
-
- memxt 徽标 -

memxt 🧠

-

经基准测试验证的最快开源 AI 记忆系统。

-

为 Claude Code 和编码智能体提供本地长期记忆——基于 SQLite + 嵌入向量的 MCP 服务器,100% 在你的设备上运行。你的智能体终于能记住昨天了。

-

- memxt Star 数 - memxt 文档 -

-
- ---- - -
- -
- -如果你读到了这里,你就是我们的同路人。
-**[⭐ 给这个仓库点个 Star](https://github.com/debpalash/VoiceStudio)**,让更多人能找到它。
-**[💬 加入 Discord](https://discord.gg/bzQavDfVV9)**,分享你的作品。
-**[❤️ 支持开发](https://ko-fi.com/debpalash)**——资助让 VoiceStudio 持续发布的 AI 智能体账单。 - -
- - - - - - Star 历史 - - -
+应用采用 [AGPL-3.0](LICENSE) 许可。模型遵循各自的许可,商用前请确认其条款。克隆声音前须取得本人许可。详见[许可说明](LICENSE-NOTICE.md)。 diff --git a/docs/integration-directory.md b/docs/integration-directory.md new file mode 100644 index 00000000..00b4f851 --- /dev/null +++ b/docs/integration-directory.md @@ -0,0 +1,20 @@ +# Integration directory + +Directory entries are illustrative, not paid sponsors, endorsements, or verified VoiceStudio integrations. Icons are bundled locally so viewing the catalog sends no logo requests to providers. Brand marks belong to their respective owners. + +| Company | Official source | Icon source | +|---|---|---| +| Twilio | [Website](https://www.twilio.com) | Bundled site icon | +| Plivo | [Website](https://www.plivo.com) | Bundled site icon | +| Telnyx | [Website](https://telnyx.com) | Bundled site icon | +| n8n | [Website](https://n8n.io) | Bundled site icon | +| Zapier | [Website](https://zapier.com) | Bundled site icon | +| Make | [Website](https://www.make.com) | Bundled generic mark | +| GitHub | [Website](https://github.com) | Bundled site icon | +| GitHub Container Registry | [Website](https://ghcr.io) | Bundled GitHub icon | +| Docker | [Website](https://www.docker.com) | Bundled site icon | +| Model Context Protocol | [Website](https://modelcontextprotocol.io) | Bundled site icon | +| OpenAI Agents | [Guide](https://platform.openai.com/docs/guides/agents) | Bundled local mark | +| Claude Code | [Guide](https://docs.anthropic.com/en/docs/claude-code) | Bundled site icon | +| Codex CLI | [Repository](https://github.com/openai/codex) | Bundled local mark | +| VoiceStudio API | [Repository](https://github.com/debpalash/VoiceStudio) | Bundled local mark | diff --git a/docs/media/electron/README.md b/docs/media/electron/README.md new file mode 100644 index 00000000..c7e402d4 --- /dev/null +++ b/docs/media/electron/README.md @@ -0,0 +1,19 @@ +# README media + +Captured from the Electron renderer on Linux, September 16, 2026. These images show the development branch, not a claim about a published release. The browser capture runs the same renderer as Electron; native window decorations are excluded. + +Only the bundled demo voice appears. Personal profiles, history, and projects are filtered from the capture context, and API mutations are blocked. The normal app and its local storage are left alone. + +With the Electron development server running: + +```bash +CHROMIUM_PATH=/usr/bin/chromium node scripts/capture-readme-electron.mjs +``` + +The script writes PNG screenshots here and prints the temporary WebM path. Convert that recording to the main GIF (replace `recording.webm` with that path): + +```bash +ffmpeg -y -ss 1 -i recording.webm -vf 'fps=8,scale=1120:-1:flags=lanczos,split[s0][s1];[s0]palettegen=stats_mode=diff[p];[s1][p]paletteuse=dither=bayer:bayer_scale=3' -loop 0 docs/media/electron/voicestudio.gif +``` + +The README uses the GIF plus the cloning and dubbing screenshots. Design and model screenshots are captured as companion stills. The official logo and repository badges retain their existing assets. diff --git a/docs/media/electron/dubbing.png b/docs/media/electron/dubbing.png new file mode 100644 index 00000000..d1745a34 Binary files /dev/null and b/docs/media/electron/dubbing.png differ diff --git a/docs/media/electron/models.png b/docs/media/electron/models.png new file mode 100644 index 00000000..5eb259fa Binary files /dev/null and b/docs/media/electron/models.png differ diff --git a/docs/media/electron/voice-cloning.png b/docs/media/electron/voice-cloning.png new file mode 100644 index 00000000..2021a657 Binary files /dev/null and b/docs/media/electron/voice-cloning.png differ diff --git a/docs/media/electron/voice-design.png b/docs/media/electron/voice-design.png new file mode 100644 index 00000000..eae7328d Binary files /dev/null and b/docs/media/electron/voice-design.png differ diff --git a/docs/media/electron/voicestudio.gif b/docs/media/electron/voicestudio.gif new file mode 100644 index 00000000..3f0daed8 Binary files /dev/null and b/docs/media/electron/voicestudio.gif differ diff --git a/docs/support-page.md b/docs/support-page.md new file mode 100644 index 00000000..1e2ed0b4 --- /dev/null +++ b/docs/support-page.md @@ -0,0 +1,19 @@ +# Support page + +The support page puts the monthly development goal, donation amounts, and Ko-fi / PayPal links first. Selecting an amount carries it into PayPal; Ko-fi lets you choose the amount on its own page. No checkout opens until you choose a provider. + +Star and community links offer other ways to help. Sponsors remain visible. The Electron page uses the official VoiceStudio logo, a single donation panel, visible sponsor and Pro cards, and an icon grid for contact channels. No accordion hides those actions. Controls support keyboard navigation, and decorative interaction animations respect reduced-motion preferences. + +Workspace headers link to Support immediately before Search. A sponsor footer sits below each workspace content area, outside its scrolling editor and above the agent dock. It reads the shared sponsor roster, uses themed hover/focus tooltips, and opens sponsor links in the system browser. With an empty roster, one combined “Your logo here” booking tile demonstrates the placement and opens the sponsorship message form. It prepares a mailto draft to partner@voicestudio.sh in the default email app, or copies the address; it never sends email itself. + +The sponsor-bar remove control opens the Free vs Pro comparison on Support. The comparison lists the proposed Pro benefits: no telemetry, a hideable sponsor bar, a Pro badge, and advanced tools. Activation verification and the specific advanced-tool list are not configured; no Pro entitlement is inferred from donations or local preferences. + +Sponsor tiles form a left-aligned, horizontally scrolling row with 1px gaps. The rightmost combined logo-plus tile opens the booking form. + +Logo hover cards match their trigger tile width and grow vertically to fit their contents. + +The footer chevron opens a searchable sponsor catalog above the strip. Cards show logos, names, tiers, and destination links; a booking card opens the email form. The panel scrolls within 60% of the viewport. Escape or the close button collapses it and returns focus to the chevron. + +The expanded catalog is labeled Integrations. Entries from the sponsored roster display a Featured badge in their catalog card and hover card; the empty booking preview does not. + +Ten voice-AI company examples populate the catalog and compact strip using locally bundled official icons. They carry a Directory example label, not Featured. Capabilities and source links are recorded in [the directory notes](integration-directory.md). diff --git a/electron/src/renderer/index.html b/electron/src/renderer/index.html index 2b4435a7..a4233b71 100644 --- a/electron/src/renderer/index.html +++ b/electron/src/renderer/index.html @@ -5,7 +5,7 @@ diff --git a/electron/src/renderer/src/components/app-shell/agent-dock-frame.tsx b/electron/src/renderer/src/components/app-shell/agent-dock-frame.tsx index ecca92a5..5ae50c6c 100644 --- a/electron/src/renderer/src/components/app-shell/agent-dock-frame.tsx +++ b/electron/src/renderer/src/components/app-shell/agent-dock-frame.tsx @@ -6,6 +6,6 @@ export function AgentDockFrame({ label, expanded = true, children }: { }) { return
{children}
; } diff --git a/electron/src/renderer/src/components/app-shell/app-shell.tsx b/electron/src/renderer/src/components/app-shell/app-shell.tsx index 8af65cac..e5f5106c 100644 --- a/electron/src/renderer/src/components/app-shell/app-shell.tsx +++ b/electron/src/renderer/src/components/app-shell/app-shell.tsx @@ -1,3 +1,4 @@ +import { SponsorFooter } from './sponsor-footer'; import { WorkspaceSidebar } from './workspace-sidebar'; import { CommandPalette } from '@/components/command-palette'; import { Outlet, useRouterState } from '@tanstack/react-router'; @@ -34,6 +35,7 @@ export function AppShell() {
+ diff --git a/electron/src/renderer/src/components/app-shell/repair-agent-dock.tsx b/electron/src/renderer/src/components/app-shell/repair-agent-dock.tsx index e75844e9..d1baabfb 100644 --- a/electron/src/renderer/src/components/app-shell/repair-agent-dock.tsx +++ b/electron/src/renderer/src/components/app-shell/repair-agent-dock.tsx @@ -282,7 +282,7 @@ export function RepairAgentDock() { return ( -
+

{t('repairAgent.title')}

@@ -347,7 +347,7 @@ export function RepairAgentDock() { ref={terminal} role="log" aria-live="polite" - className="studio-scrollbar min-h-0 min-w-0 flex-1 overflow-auto whitespace-pre-wrap break-words bg-[var(--app-theme-terminal-background,var(--background))] p-4 font-mono text-xs leading-5 text-[var(--app-theme-terminal-foreground,var(--foreground))]" + className="studio-scrollbar min-h-0 min-w-0 flex-1 overflow-auto whitespace-pre-wrap break-words bg-[var(--app-theme-terminal-background,var(--background))] p-3 font-mono text-xs leading-5 text-[var(--app-theme-terminal-foreground,var(--foreground))]" > {output} @@ -355,7 +355,7 @@ export function RepairAgentDock() {

@@ -369,7 +369,7 @@ export function RepairAgentDock() {

)} -
+
{!workspaceAvailable && !appOperation && ( +
+ +
+ {visibleSponsors.map((sponsor) => ( + { + const bridge = getBridge(); + if (!bridge) return; + event.preventDefault(); + setFailed(false); + void bridge.files.openExternal(sponsor.url).catch(() => setFailed(true)); + }} + > +
+ { + event.currentTarget.style.display = 'none'; + }} + /> +
+
+

{sponsor.name}

+ + {t( + sponsor.featured + ? 'integrationCatalog.featured' + : 'directoryExamples.example', + )} + +
+ {SPONSOR_TIERS.includes(sponsor.tier) && ( + {t('support.sponsors_tier_' + sponsor.tier)} + )} + {sponsor.detailKeys.length > 0 && ( +

{sponsor.detailKeys.map((key) => t(key)).join(' · ')}

+ )} +

{sponsor.url}

+
+ ))} + {entries.length > 0 && visibleSponsors.length === 0 && ( +

+ {t('common.no_matches')} +

+ )} + +
+ + )} +
+ +
{ + const bounds = event.currentTarget.getBoundingClientRect(); + const edge = Math.min(72, bounds.width * 0.18); + edgeScrollRef.current = + event.clientX < bounds.left + edge ? -5 : event.clientX > bounds.right - edge ? 5 : 0; + }} + onMouseLeave={() => { + edgeScrollRef.current = 0; + }} + > + {entries.map((sponsor) => ( + + { + const bridge = getBridge(); + if (!bridge) return; + event.preventDefault(); + setFailed(false); + void bridge.files.openExternal(sponsor.url).catch(() => setFailed(true)); + }} + /> + } + > + { + event.currentTarget.style.display = 'none'; + }} + /> + {sponsor.name} + + + + + {sponsor.name} + + + {t( + sponsor.featured ? 'integrationCatalog.featured' : 'directoryExamples.example', + )} + + {SPONSOR_TIERS.includes(sponsor.tier) && ( + + {t('support.sponsors_tier_' + sponsor.tier)} + + )} + + {sponsor.detailKeys.length + ? sponsor.detailKeys.map((key) => t(key)).join(' · ') + : t('integrationCatalog.description')} + + + {sponsor.url} + + + + ))} +
+ {failed && ( + + {t('common.error')} + + )} + + setInquiryOpen(true)} + className="sponsor-book-tile sponsor-book-tile--combined" + aria-label={t('sponsorSlot.book')} + /> + } + > + + + {t('sponsorSlot.footer_brand')} + {t('sponsorSlot.footer_book')} + + + + + + {t('sponsorSlot.title')} + + {t('sponsorSlot.description')} + + + {t('support.sponsors_perk')} + + + {t('sponsorSlot.preview_detail')} + + + + + + + void navigate({ to: '/settings/support', search: { compare: true } }) + } + /> + } + > + + + {t('supportPlans.remove')} + + + +
+ + ); +} diff --git a/electron/src/renderer/src/components/app-shell/sponsor-inquiry.css b/electron/src/renderer/src/components/app-shell/sponsor-inquiry.css new file mode 100644 index 00000000..6ee787dc --- /dev/null +++ b/electron/src/renderer/src/components/app-shell/sponsor-inquiry.css @@ -0,0 +1,84 @@ +.sponsor-inquiry-dialog { + display: flex; + flex-direction: column; + min-height: 0; + border: 1px solid var(--sidebar-border); + background: + radial-gradient(circle at 8% 0%, color-mix(in srgb, var(--primary) 12%, transparent), transparent 34%), + var(--sidebar); + color: var(--sidebar-foreground); + box-shadow: + 0 24px 70px rgb(0 0 0 / 38%), + inset 0 1px 0 rgb(255 255 255 / 6%); +} +.sponsor-inquiry-tabs { + border: 1px solid var(--sidebar-border); + background: color-mix(in srgb, var(--sidebar-foreground) 4%, var(--sidebar)); +} +.sponsor-inquiry-hero-icon { + display: grid; + place-items: center; + width: 42px; + height: 42px; + flex: 0 0 42px; + border: 1px solid color-mix(in srgb, var(--primary) 35%, var(--sidebar-border)); + border-radius: 11px; + background: color-mix(in srgb, var(--primary) 14%, var(--sidebar)); + color: var(--primary); + box-shadow: inset 0 1px 0 rgb(255 255 255 / 8%); +} +.sponsor-inquiry-hero-icon svg { width: 21px; height: 21px; } +.sponsor-inquiry-perks { + display: grid; + grid-template-columns: repeat(3, minmax(0, 1fr)); + gap: 8px 18px; +} +.sponsor-inquiry-perks span { + display: flex; + min-width: 0; + align-items: center; + gap: 3px; + color: var(--muted-foreground); +} +.sponsor-inquiry-perks span + span { + border-left: 0; + padding-left: 0; +} +.sponsor-inquiry-perks svg { width: 13px; height: 13px; flex: 0 0 auto; color: var(--primary); } +.sponsor-inquiry-perks small { font-size: 10px; line-height: 1.2; } +.sponsor-inquiry-tab { + border-radius: 8px; + color: var(--muted-foreground); +} +.sponsor-inquiry-tab[aria-selected='true'] { + border-color: color-mix(in srgb, var(--sidebar-border) 80%, var(--primary)); + background: linear-gradient(180deg, color-mix(in srgb, var(--primary) 18%, var(--sidebar-accent)), var(--sidebar-accent)); + color: var(--sidebar-foreground); + box-shadow: + inset 0 1px 0 rgb(255 255 255 / 8%), + 0 4px 12px rgb(0 0 0 / 12%); +} +.sponsor-inquiry-tab:hover { + color: var(--sidebar-foreground); +} +.sponsor-inquiry-form-link { + width: 32px; + color: var(--muted-foreground); +} +.sponsor-inquiry-form-link svg { width: 14px; height: 14px; } +.sponsor-inquiry-form-link:hover { + color: var(--sidebar-foreground); + background: color-mix(in srgb, var(--primary) 10%, transparent); +} +.sponsor-inquiry-panel { + min-height: 0; + border-color: var(--sidebar-border); + background: color-mix(in srgb, var(--sidebar-foreground) 2%, var(--sidebar)); +} +.sponsor-inquiry-panel > button[type='submit'] { + margin-top: auto; + flex: 0 0 auto; +} +@media (max-width: 560px) { + .sponsor-inquiry-perks { grid-template-columns: repeat(2, minmax(0, 1fr)); } +} diff --git a/electron/src/renderer/src/components/app-shell/sponsor-inquiry.tsx b/electron/src/renderer/src/components/app-shell/sponsor-inquiry.tsx new file mode 100644 index 00000000..e04ba529 --- /dev/null +++ b/electron/src/renderer/src/components/app-shell/sponsor-inquiry.tsx @@ -0,0 +1,236 @@ +import { useState } from 'react'; +import { + BlocksIcon, + BookOpenIcon, + BadgeCheckIcon, + CopyIcon, + DownloadIcon, + EyeIcon, + ExternalLinkIcon, + MailIcon, + ShieldCheckIcon, + PinIcon, + XIcon, +} from 'lucide-react'; +import { useTranslation } from 'react-i18next'; +import { + Dialog, + DialogContent, + DialogTitle, + DialogDescription, + DialogClose, +} from '@/components/ui/dialog'; +import { Button } from '@/components/ui/button'; +import { getBridge } from '@/components/bridge'; +import './sponsor-inquiry.css'; + +export const PARTNER_EMAIL = 'partner@voicestudio.sh'; +export const SPONSOR_FORM_URL = 'https://forms.gle/2PYCvd39hbwijzX37'; +export function sponsorMailto(subject: string, message: string) { + return `mailto:${PARTNER_EMAIL}?subject=${encodeURIComponent(subject)}&body=${encodeURIComponent(message)}`; +} + +export function SponsorInquiry({ + open, + onOpenChange, +}: { + open: boolean; + onOpenChange: (open: boolean) => void; +}) { + const { t } = useTranslation(); + const emailTemplate = t('sponsorSlot.email_template'); + const [message, setMessage] = useState(emailTemplate); + const [copied, setCopied] = useState(false); + const [failed, setFailed] = useState(false); + const [busy, setBusy] = useState(false); + const [mode, setMode] = useState<'form' | 'email'>('form'); + return ( + { + setCopied(false); + setFailed(false); + onOpenChange(value); + }} + > + + + } + > + +
+ +
+ {t('sponsorSlot.title')} + {t('sponsorSlot.description')} +
+
+
+ + + + + + + + + + + + +
+
+ + + +
+ {mode === 'form' ? ( +