Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
3.6 KiB
VoiceStudio — OpenAI-Compatible Remote ASR
Point transcription at any server exposing an OpenAI-compatible
POST /v1/audio/transcriptions endpoint — LM Studio or a llama.cpp-style
local server, a self-hosted Qwen3-ASR/FunASR/SenseVoice box on your network,
Groq, or OpenAI's own Whisper API. Unlike every other ASR engine, this one
runs no model locally: it's a pure network client, so it needs no install
and claims no GPU.
Setup
Everything lives on one screen — Settings → Engines, ASR tab:
- The OpenAI-compatible (remote server) row shows as unavailable until a server is configured. The config panel appears below the engine list while the ASR tab is selected.
- Set Server URL to your server's base URL (see the examples below).
- Set Model to whatever your server expects (
whisper-1for OpenAI's API; check your server's docs or its/v1/modelslisting otherwise). - API key is optional — many self-hosted servers accept requests without one. Set it if your server requires auth. The key is stored encrypted on your machine and is never displayed or sent anywhere except the server you configured.
- Click Test connection. This saves the fields, then sends one tiny
GET /modelsrequest to the server — no audio is uploaded, nothing is transcribed. You'll see the round-trip latency on success (plus whether your configured model is in the server's list), or the exact failure (unreachable / timeout / rejected key / HTTP status) if not. - Click Use on the engine's row to make it the active ASR engine — the
same picker every engine family has. Power users can pin it instead with
OMNIVOICE_ASR_BACKEND=openai-compat-asrbefore launching; the env var always wins over the Settings pick.
Config changes apply on the next transcription — no restart needed. The engine is never active by default: VoiceStudio's ASR auto-detect only ever picks local engines, and the app works fully with this engine unconfigured.
Examples
| Server | Server URL | Model | API key |
|---|---|---|---|
| LM Studio (local) | http://localhost:1234/v1 |
the model name shown in LM Studio | none |
| llama.cpp / whisper.cpp server (local) | http://localhost:8080/v1 |
whatever the server loads (often ignored) | none |
| speaches / faster-whisper-server (local) | http://localhost:8000/v1 |
e.g. Systran/faster-whisper-large-v3 |
none |
| Self-hosted Qwen3-ASR / FunASR (LAN box) | http://<host>:8000/v1 |
your deployment's model id | if you enabled auth |
| Groq | https://api.groq.com/openai/v1 |
whisper-large-v3 |
required |
| OpenAI | https://api.openai.com/v1 |
whisper-1 |
required |
Local servers vary in which endpoints they implement — if Test connection reports the server is reachable but doesn't list models, transcription may still work; run a small dictation or dub-transcribe to confirm.
Response format
The backend prefers response_format=verbose_json for real per-segment
timestamps (OpenAI's API and most compatible servers support it) and falls
back to plain text automatically if your server rejects that format. Neither
path returns word-level timestamps — that's not part of this API.
Privacy note
Unlike every other ASR engine in VoiceStudio, audio sent through this backend leaves your machine — to whatever server you configured, and nowhere else. If that's a self-hosted server on your own network, nothing leaves your control; if it's a third-party API (Groq, OpenAI's, or someone else's), review their data handling before sending anything sensitive.