Files
VoiceStudio/docs/engines/supertonic3.md
T
Palash DebnathandClaude Fable 5 030d5ea01f docs(engines): a guide for every engine + index; fix two engine-metadata bugs (#1556)
* docs(engines): a guide for every engine + index; fix two engine-metadata bugs

21 new pages under docs/engines/ (10 TTS, 10 ASR, index README) — every
registered engine now has one: what it's for, platform support, model env
vars, quirks with issue refs. Linked from both READMEs' engine sections.

Code fixes found while verifying facts against the registries:
- KittenTTS docstring claimed default voice 'Jasper'; the code default is
  expr-voice-2-f
- the isolated-ASR sidecar read only ASR_MODEL_FW while the download
  preflight read ASR_MODEL_FASTER — set one and the other quietly used a
  different model; both now resolve ASR_MODEL_FW-override → ASR_MODEL_FASTER
- moonshine's install hint named 'useful-moonshine', a package the backend
  never imports; now moonshine-onnx / moonshine-voice

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(changelog): entries for the engine guides + sidecar model fix (#1556)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(engines): second-harvest fixes — TLS guidance, matrix/code alignment, CN counts

- README matrix aligned to gpu_compat (the code is the source of truth):
  CosyVoice macOS is CPU not MPS, IndexTTS and GGUF gain their real
  CUDA/CPU/MPS cells
- gpt-sovits guide: prefer https/tunnel for non-loopback servers, plaintext
  warning; first-use download guidance on both OmniVoice pages
- preflight empty-env fallback matches the sidecar (ASR_MODEL_FASTER='' no
  longer resolves a different repo)
- nano installs via uv pip; kitten log level wording; index links install
  guides incl. the Gatekeeper step; README_CN engine counts 16/11

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(readme-cn): the all-engines-local claim now excludes the remote client

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 03:55:59 +00:00

2.8 KiB
Raw Blame History

VoiceStudio — Supertonic-3 Engine

Supertonic-3 (Supertone Inc.) is a ~99M-parameter ONNX TTS engine covering 31 languages with 7 preset voices at native 44.1 kHz. It is CPU-only by design — pure ONNX Runtime on the CPU execution provider, with no CUDA or MPS path in the upstream SDK — and runs in its own sidecar process so crashes and cold init never block the rest of VoiceStudio.

When to pick it

  • Broad language coverage on machines with no usable GPU.
  • Preset-voice narration at a higher sample rate than the default engine.

Setup

  1. Install the optional dependency into VoiceStudio's environment:

    uv sync --extra supertonic
    

    (Or enable it from Model Catalogue → Engines, which installs the pinned supertonic wheel for you.)

  2. Accept the license in-app. First use is gated behind an explicit acceptance dialog: the inference SDK is MIT, but the model weights are OpenRAIL-M, which carries use restrictions. The engine stays unavailable until you review and accept in Model Catalogue → Engines → Supertonic-3.

  3. Select the engine via Model Catalogue → Engines or OMNIVOICE_TTS_BACKEND=supertonic3.

The first synthesis cold-downloads ~400 MB of model weights, pinned to an exact HuggingFace revision SHA so the bytes match what the SDK was validated against. See downloading-models.md.

Voices

Seven preset voices are surfaced: M1 (default), M3, M4, M5, F3, F4, F5. The SDK itself accepts the full M1M5 / F1F5 set if a caller passes one explicitly; unknown ids fall back to the default with a log line.

Behaviour notes

  • Output is 44.1 kHz mono.
  • Runs as a long-lived sidecar in the parent Python environment (its dependencies — onnxruntime, numpy, soundfile — already match VoiceStudio's pins); subsequent calls reuse the warm ONNX session.
  • speed is clamped to 0.72.0; quality steps clamp to 512.
  • Language is an ISO 639-1 code; Auto engages the SDK's multilingual fallback.

Known limits

  • No cloning and no voice design — preset voices only. Dub/batch jobs that need cloning won't select it.
  • CPU-only: hardware acceleration is a property of the upstream SDK, not a VoiceStudio limitation.
  • OpenRAIL-M weights are not covered by VoiceStudio's blanket commercial-use statement — review the model license terms in the acceptance dialog.

Troubleshooting

  • "supertonic package not installed": run the uv sync above or enable from the Model Catalogue.
  • "license not accepted": open Model Catalogue → Engines → Supertonic-3 and accept.
  • Other issues: install/troubleshooting.md.

See also: benchmarks.md, languages.md, expressive-speech.md, disk usage.