Files
VoiceStudio/docs/engines/audio-cpp.md
T

5.1 KiB

VoiceStudio — audio.cpp Engine (Breeze-TTS-2)

audio.cpp is a pure-C++ ggml audio inference framework with prebuilt binaries for Windows, macOS, and Linux and no Python dependency. VoiceStudio's initial integration runs it on CPU. VoiceStudio drives its audiocpp_server over loopback HTTP — v1 serves the breeze_tts family: Breeze-TTS-2 (BreezeBlue, 3B params, English + Chinese, voice clone + voice design + voice direction, 24 kHz).

Opt-in, and never a default. Select audiocpp explicitly in Model Catalogue → Engines (or OMNIVOICE_TTS_BACKEND=audiocpp).

License — read before enabling

  • audio.cpp code: Apache-2.0.
  • Breeze-TTS-2 weights (upstream BreezeBlue/Breeze-TTS-2 and the audio-cpp/audio.cpp-gguf GGUF repack): research and non-commercial use only under the BreezeBlue Research and Non-Commercial License. Self-hosted outputs inherit the restriction; a BreezeBlue paid subscription covers hosted-platform outputs only, not this engine.

Platform support

Host Binary Compute
Windows x64 portable CPU prebuilt CPU
Linux x64 CPU prebuilt CPU
macOS arm64 / x64 upstream macOS prebuilt CPU
Linux aarch64 none upstream unavailable in v1

GPU acceleration is outside this first integration. It will be enabled only after backend selection and physical-device routing are verified per platform. The Q8_0 GGUF is approximately 4.73 GiB, plus CPU runtime memory.

Install

  1. Download the v0.7.2 prebuilt for your platform from audio.cpp releases (use the CPU archive on Windows or Linux) and extract it. VoiceStudio does not download executable code for this engine. The Linux archive does not preserve the executable bit, so run chmod +x audiocpp_server after extracting it.

    Verify the archive before extracting it. The pinned SHA-256 checksums are:

    Archive SHA-256
    audio-v0.7.2-bin-windows-x64-cpu-portable.zip 0b1f4bd78c5226ee3fa0eb24d95d603a429439cdf5dab45872d44a87412dd8c1
    audio-v0.7.2-bin-ubuntu-x64-cpu.tar.gz 6f5e43dd7b80e8ddf688ef84b411fadcd1f934d2c83963178bc4e2d9c4f07736
    audio-v0.7.2-bin-macos-arm64-metal.tar.gz c01e4f82971bedbe341697e63a9cebd5a5d1f72d5a9bcb51a3191f95ddab7a95
    audio-v0.7.2-bin-macos-x64-metal.tar.gz 3862270f33439077225324169313f727064f727305b54d8ce920244d75ddcc24

    Run sha256sum <archive> on Linux, shasum -a 256 <archive> on macOS, or Get-FileHash <archive> -Algorithm SHA256 in PowerShell and compare the complete result with the table.

  2. Set OMNIVOICE_AUDIOCPP_BIN to the audiocpp_server binary (audiocpp_server.exe on Windows):

    # macOS / Linux
    echo 'export OMNIVOICE_AUDIOCPP_BIN=$HOME/apps/audio.cpp/audiocpp_server' >> ~/.zshrc
    source ~/.zshrc
    

    Alternatively set OMNIVOICE_AUDIOCPP_DIR to the directory containing it.

  3. Restart VoiceStudio. The ~4.73 GiB breeze-tts-2-q8_0.gguf downloads from audio-cpp/audio.cpp-gguf (not gated) into the shared HF cache on first generate — resumable, hash-verified.

  4. Pick audiocpp in Model Catalogue → Engines. The server starts lazily on first generate (server.json + server.log live under the app data audiocpp/ directory).

Voice modes

All three go through the one speech endpoint — reference presence selects:

  • Clone: ref_audio + ref_text (exact transcript, as upstream).
  • Direction: ref_audio + ref_text + instruct (e.g. "Speak slowly with a restrained, serious tone").
  • Design: description (or instruct) with no ref_audio (e.g. "A warm, thoughtful young woman…"). Upstream strengthens instruction-following with guidance_scale ≈ 4.

Optional env knobs

Variable Default Purpose
OMNIVOICE_AUDIOCPP_BIN — Absolute path to audiocpp_server.
OMNIVOICE_AUDIOCPP_DIR — Directory containing audiocpp_server.
OMNIVOICE_AUDIOCPP_MODEL pinned auto-download GGUF file or directory override.
OMNIVOICE_AUDIOCPP_PACKAGE breeze-tts-2-q8_0.gguf Package filename (…-bf16.gguf for full precision).
OMNIVOICE_AUDIOCPP_PORT 17860 Loopback port.

Common errors

audiocpp_server not found ...

The binary isn't installed. Follow Install — the message carries the exact release URL and SHA for your platform.

audiocpp_server exited during startup ...

The managed loopback port may be taken. Check server.log next to server.json in the app data audiocpp/ directory, or set a different OMNIVOICE_AUDIOCPP_PORT and restart VoiceStudio.

Breeze-TTS-2 package ... missing after download

The upstream audio.cpp-gguf layout changed. File an issue with the package listing — the allow-list in bootstrap.py needs updating.


audio.cpp runs as a managed native server (no Python venv, no transformers conflict). Only the downloaded GGUF counts toward sidecar disk usage.