Files
VoiceStudio/docs/engines/supertonic3.md
T
Palash Debnath ae68dd5bfd fix(engines): a half-finished install is repaired, not reported installed
An engine with no weights download counted as installed once its venv
interpreter existed, so a dependency install that died halfway made the
next attempt answer already_installed and the engine failed at its first
import. The import probe now writes a completion marker, and a fresh
dependency step removes the old one. IndexTTS keeps its weights check, so
no existing install is asked to reinstall.

The MOSS bootstrap no longer blames a non-CUDA host for an install
failure; the index is always supplied, and uv's error says what failed.
The PocketTTS and Supertonic guides now say where the sidecar runs, and
the three repository-engine guides say what to do if the first weight
download outlasts the compute-time budget.
2026-09-10 08:12:17 -07:00

3.1 KiB
Raw Blame History

VoiceStudio — Supertonic-3 Engine

Supertonic-3 (Supertone Inc.) is a ~99M-parameter ONNX TTS engine covering 31 languages with 7 preset voices at native 44.1 kHz. It is CPU-only by design — pure ONNX Runtime on the CPU execution provider, with no CUDA or MPS path in the upstream SDK — and runs in its own sidecar process so crashes and cold init never block the rest of VoiceStudio.

When to pick it

  • Broad language coverage on machines with no usable GPU.
  • Preset-voice narration at a higher sample rate than the default engine.

Setup

  1. Install the optional dependency into VoiceStudio's environment:

    uv sync --extra supertonic
    

    Or click Install in Model Catalogue → Engines → Supertonic-3. That installs the same pinned wheel into the engine's own Python environment under VoiceStudio's data directory. Nothing it installs touches VoiceStudio itself or any other engine, and Uninstall in the same row removes only that folder. An install made with uv sync keeps working as it is.

  2. Accept the license in-app. First use is gated behind an explicit acceptance dialog: the inference SDK is MIT, but the model weights are OpenRAIL-M, which carries use restrictions. The engine stays unavailable until you review and accept in Model Catalogue → Engines → Supertonic-3.

  3. Select the engine via Model Catalogue → Engines or OMNIVOICE_TTS_BACKEND=supertonic3.

The first synthesis cold-downloads ~400 MB of model weights, pinned to an exact HuggingFace revision SHA so the bytes match what the SDK was validated against. See downloading-models.md.

Voices

Seven preset voices are surfaced: M1 (default), M3, M4, M5, F3, F4, F5. The SDK itself accepts the full M1M5 / F1F5 set if a caller passes one explicitly; unknown ids fall back to the default with a log line.

Behaviour notes

  • Output is 44.1 kHz mono.
  • Runs as a long-lived sidecar: from its own environment after a one-click install, otherwise from VoiceStudio's (where uv sync --extra supertonic puts it). Subsequent calls reuse the warm ONNX session.
  • speed is clamped to 0.72.0; quality steps clamp to 512.
  • Language is an ISO 639-1 code; Auto engages the SDK's multilingual fallback.

Known limits

  • No cloning and no voice design — preset voices only. Dub/batch jobs that need cloning won't select it.
  • CPU-only: hardware acceleration is a property of the upstream SDK, not a VoiceStudio limitation.
  • OpenRAIL-M weights are not covered by VoiceStudio's blanket commercial-use statement — review the model license terms in the acceptance dialog.

Troubleshooting

  • "supertonic package not installed": run the uv sync above or enable from the Model Catalogue.
  • "license not accepted": open Model Catalogue → Engines → Supertonic-3 and accept.
  • Other issues: install/troubleshooting.md.

See also: benchmarks.md, languages.md, expressive-speech.md, disk usage.