7.3 KiB
VoiceStudio — audio.cpp Engine (Breeze-TTS-2)
audio.cpp is a pure-C++ ggml audio
inference framework with prebuilt binaries for Windows, macOS, and Linux and
no Python dependency. VoiceStudio discovers the compute providers compiled
into the installed binary and selects the best available device. VoiceStudio drives its
audiocpp_server over loopback HTTP — v1 serves the breeze_tts
family: Breeze-TTS-2 (BreezeBlue, 3B params, English + Chinese, voice
clone + voice design + voice direction, 24 kHz).
Opt-in, and never a default. Select
audiocppexplicitly in Model Catalogue → Engines (orOMNIVOICE_TTS_BACKEND=audiocpp).
License — read before enabling
- audio.cpp code: Apache-2.0.
- Breeze-TTS-2 weights (upstream
BreezeBlue/Breeze-TTS-2and theaudio-cpp/audio.cpp-ggufGGUF repack): research and non-commercial use only under the BreezeBlue Research and Non-Commercial License. Self-hosted outputs inherit the restriction; a BreezeBlue paid subscription covers hosted-platform outputs only, not this engine.
Platform support
| Host | Binary | Compute |
|---|---|---|
| Windows x64 | CPU, Vulkan, CUDA 12.4/13.3 prebuilts | CPU, Vulkan, CUDA |
| Linux x64 | CPU and Vulkan prebuilts | CPU, Vulkan; CUDA/ROCm from a self-build |
| macOS arm64 | upstream Metal prebuilt | Metal + CPU |
| macOS x64 | upstream metal archive (Metal disabled by upstream) |
CPU |
| Linux aarch64 | none upstream | unavailable in v1 |
VoiceStudio runs audiocpp_server --list-devices once per installed binary,
prefers CUDA, ROCm, Metal, then a discrete Vulkan GPU, and keeps CPU as a safe
fallback. Device numbers are local to each backend registry. The Q8_0 GGUF is
approximately 4.73 GiB, plus runtime memory.
Allow about 6 GB of dedicated VRAM for CUDA, HIP, or a discrete GPU exposed through Vulkan. Metal and Vulkan integrated GPUs use unified-memory handling instead of that dedicated-VRAM floor.
Install
-
Download the v0.7.2 prebuilt for your platform from audio.cpp releases and extract it. Use the Vulkan archive on Windows or Linux for broad GPU support, the CPU archive when Vulkan is unavailable, or the matching CUDA archive on Windows for NVIDIA. A Windows CUDA install needs both the
bin-…-cuda…and matchingcudart-…-cuda…archives extracted into the same directory, as required by upstream. VoiceStudio does not download executable code for this engine. Linux archives do not preserve the executable bit, so runchmod +x audiocpp_serverafter extracting one.Verify the archive before extracting it. The pinned SHA-256 checksums are:
Archive SHA-256 audio-v0.7.2-bin-windows-x64-cpu-portable.zip0b1f4bd78c5226ee3fa0eb24d95d603a429439cdf5dab45872d44a87412dd8c1audio-v0.7.2-bin-windows-x64-vulkan.zip15b8232eae740e21e507d87f827a89966de9451b085a45932d9e214e032962c1audio-v0.7.2-bin-windows-x64-cuda12.4.zip06c426095008022a2984ff1c75de4c9fab463c4201c0ff0a5dc4e14043f52326audio-v0.7.2-cudart-windows-x64-cuda12.4.zip7115be4d462817ad293f7932a8ac436d51023128e6728af09bba92a85593f393audio-v0.7.2-bin-windows-x64-cuda13.3.zipf975fec52745807b8c787e826c110acc3424a45156092455e1635604b68ec832audio-v0.7.2-cudart-windows-x64-cuda13.3.zip9b508f702636a9cdf3bf4dd8e75a86c20a0b87bdc39ca07e714c82f748efc1faaudio-v0.7.2-bin-ubuntu-x64-cpu.tar.gz6f5e43dd7b80e8ddf688ef84b411fadcd1f934d2c83963178bc4e2d9c4f07736audio-v0.7.2-bin-ubuntu-x64-vulkan.tar.gzfee1f978cee76453cf17f00196554bc2ee294645739538af0726a143b6a69a23audio-v0.7.2-bin-macos-arm64-metal.tar.gzc01e4f82971bedbe341697e63a9cebd5a5d1f72d5a9bcb51a3191f95ddab7a95audio-v0.7.2-bin-macos-x64-metal.tar.gz3862270f33439077225324169313f727064f727305b54d8ce920244d75ddcc24Run
sha256sum <archive>on Linux,shasum -a 256 <archive>on macOS, orGet-FileHash <archive> -Algorithm SHA256in PowerShell and compare the complete result with the table. -
Set
OMNIVOICE_AUDIOCPP_BINto theaudiocpp_serverbinary (audiocpp_server.exeon Windows):# macOS / Linux echo 'export OMNIVOICE_AUDIOCPP_BIN=$HOME/apps/audio.cpp/audiocpp_server' >> ~/.zshrc source ~/.zshrcAlternatively set
OMNIVOICE_AUDIOCPP_DIRto the directory containing it. -
Restart VoiceStudio, open Model Catalogue → Models, find Breeze-TTS-2 Q8_0 for audio.cpp, review its research/non-commercial license note, and click Install. Generation never starts this ~4.73 GiB download automatically.
-
Pick
audiocppin Model Catalogue → Engines. The server starts lazily on first generate (server.json+server.loglive under the app dataaudiocpp/directory).
Voice modes
All three go through the one speech endpoint — reference presence selects:
- Clone:
ref_audio+ref_text(exact transcript, as upstream). - Direction:
ref_audio+ref_text+instruct(e.g. "Speak slowly with a restrained, serious tone"). - Design:
description(orinstruct) with noref_audio(e.g. "A warm, thoughtful young woman…"). Upstream strengthens instruction-following withguidance_scale≈ 4.
Optional env knobs
| Variable | Default | Purpose |
|---|---|---|
OMNIVOICE_AUDIOCPP_BIN |
— | Absolute path to audiocpp_server. |
OMNIVOICE_AUDIOCPP_DIR |
— | Directory containing audiocpp_server. |
OMNIVOICE_AUDIOCPP_MODEL |
Model Catalogue cache | GGUF file or directory override. |
OMNIVOICE_AUDIOCPP_PACKAGE |
breeze-tts-2-q8_0.gguf |
Package filename (…-bf16.gguf for full precision). |
OMNIVOICE_AUDIOCPP_PORT |
17860 |
Loopback port. |
OMNIVOICE_AUDIOCPP_BACKEND |
Settings, then auto | Exact runtime: cuda, hip/rocm, vulkan, metal, or cpu. |
OMNIVOICE_AUDIOCPP_DEVICE |
best device | Backend-local non-negative device index; requires OMNIVOICE_AUDIOCPP_BACKEND. |
The audio.cpp overrides take precedence over the global Settings compute choice. A global CUDA/ROCm choice can match an NVIDIA/AMD GPU exposed through Vulkan. An unavailable explicit audio.cpp override is an error; an unavailable global preference falls back to CPU and is shown as a routing fallback.
Common errors
audiocpp_server not found ...
The binary isn't installed. Follow Install — the message carries the exact release URL and SHA for your platform.
audiocpp_server exited during startup ...
The managed loopback port may be taken. Check server.log next to
server.json in the app data audiocpp/ directory, or set a different
OMNIVOICE_AUDIOCPP_PORT and restart VoiceStudio.
Breeze-TTS-2 ... not installed or package ... not completely installed
Install the model from Model Catalogue → Models. If an interrupted install left it incomplete, use Reinstall there. If the error persists after a complete reinstall, file an issue with the package listing.
audio.cpp runs as a managed native server (no Python venv, no
transformers conflict). Only the downloaded GGUF counts toward
sidecar disk usage.