Exports failed with a 422 naming a field the current app never sends — twice, from different users. The cause was the attach handshake: if something already answers on the backend port and reports a matching version, the app adopts it and skips the source sync a normal launch performs. A version string holds steady for a whole release cycle, so a same-version process can still be running weeks-old code, and that code then serves a current UI.
The handshake now compares a fingerprint of the shipped Python sources, read from the same response as the version so a dropped probe can't masquerade as a missing field. A backend predating the mechanism is treated as stale; one that is current but started outside the app is still accepted. Refusals are logged with a greppable marker, since this class previously took two reports and a code audit to identify.
Fixes#1770. Closes the duplicate report tracked in #1792.
On Windows with a non-UTF-8 system code page, a venv path containing non-English characters (commonly a CJK username) killed the interpreter during startup, before any VoiceStudio code ran — Python reads .pth files in the active code page, and uv's editable install writes the project path there in UTF-8. The backend could never start, and the setup screen only said it had stalled.
A new or genuinely broken environment now builds at an ASCII-safe short path; an existing environment that starts cleanly is never relocated. When no safe path can be produced, the app says so before downloading rather than after. The failure message names the cause and a remedy that works for the install mode in use, since portable installs ignore the setting managed installs use.
Fixes#1783.
The compute-time error told users to raise a generation timeout that had no control anywhere in the app — the only knob was an environment variable, and on Windows the docs explicitly warn against the usual way of setting one. Both budgets are now editable in Settings under Performance & Device, persisted and applied on the next start.
Two defects found in review and fixed here rather than shipped: an explicit universal budget silently overrode a separately saved CPU budget, so the CPU row would have looked like it worked and done nothing; and a value already set in the environment shadowed the saved preference while the panel still reported success. A shadowed row now says so instead. Long-input warnings also fire on Apple Silicon, which gets the accelerated budget and was the device in one of the duplicate reports.
Fixes#1787. Closes the reports tracked in #1774 and #1778.
The panel printed every chosen value three times and held twelve control rows open before anyone touched it. Details now collapse to a single recipe line that expands on request, gender/age/pitch/style become selects, and English accent and Chinese dialect merge into one grouped field so the combination the engine rejects cannot be selected at all. Starting points show five with an overflow, the bottom-bar slider is labelled Steps, and the panel ends where its content ends.
Also fixes two races found in review: a manual pick or a cleared description now beats an in-flight describe response, and arrow-key chip navigation moves focus without resetting the design.
Exports and every other native-picker action 403'd with "Invalid or expired desktop authorization" whenever Tauri and the backend resolved different data directories — a dev backend spawned without the OMNIVOICE_* env, a custom data folder, or portable mode. Tauri now takes the directory the running backend advertises, so the two processes cannot disagree, falling back to its own resolution when the backend is unreachable. One resolver covers all six capability kinds. Also decodes JSON escapes when reading that field, which Windows paths depend on, requires an absolute path so a relative data dir cannot recreate the same split, and keeps the 403 and its log line free of filesystem paths. Fixes#1781.
Voice Design rendered EnglishAccent and ChineseDialect as two unlinked controls, so a user could select both and only learn they conflict from a 400 after a round trip. A shared exclusive-groups map now mirrors the engine's rule across every path that builds or restores instruct state: the live picker, free-text entry, saved-profile and imported-session restore, plus a message-matching backstop for a conflict arriving by any other route. Picking one clears the other with a visible reason instead of a silent reset. Fixes#1771.
The Japanese clone.cleaning status read 掃除中 (tidying up a room) instead of ノイズ除去中, which is what the step actually does: denoising the reference audio. Thanks @j30231!
Corrects 231 machine-translation defects in the Korean locale and translates all 493 previously missing keys, dropping ko to zero in the missing-key baseline. Includes terminology consistency (Cinematic, export, UI scale) and fixes copy that named the wrong control. Maintainer commits merged current main and folded in the five live-preview keys added by #1769. Thanks @j30231!
Opt-in watch folder on the batch queue: pick a directory once and new videos are auto-enqueued with the last Add-to-queue settings, with pause/stop controls and copy-in-progress protection. Files stream to the loopback backend as bytes; paths never leave the app. Also gives the batch queue a reachable UI entry point and streams multipart uploads to disk. Maintainer fix: the watched directory handle is opened with full share mode on Windows so users can rename or delete the folder while it is watched, matching macOS/Linux behaviour, with a cross-platform regression test. Thanks @mvanhorn!
Opt-in live preview for dub segments: edits debounce into a streamed /ws/tts synthesis played through the chunk player, with cancellation preserved through buffered playback. Maintainer fixes: /ws/tts added to the backend ticket allowlist (feature was dead off-loopback), handshake failures surface a toast, loopback-only plaintext refusal reverted to keep the documented remote-GPU setup working, PCM16 decode hardened. Thanks @mvanhorn!
* docs: polish README hero hierarchy
* docs: enrich README with audio samples, hardware guide, and doc links
* docs: address review findings on Docker port binding, MCP transport, and privacy
* docs: address CodeRabbit review feedback on cURL format and Colab links
* docs: align Docker quick-start with stable tag and named volume mount
* feat(audiobook): synced-lyrics player
Replace the bare <audio controls> in the audiobook result with a player
that renders the chapter text and highlights the word under
audio.currentTime, karaoke-style. Timing reuses what the render stream
already emits — per-chapter duration_s on the chapter SSE events — with
words even-split inside each chapter (the karaoke burn-in's old-job
fallback, ported from services/karaoke_ass.py); after a reload the whole
book even-splits over the file's own duration. No ASR pass, no new
backend surface, nothing leaves the machine. Download keeps going
through the Tauri-safe downloadMedia util.
New pure helper utils/audiobookLyrics.js mirrors the longform parser's
chapter drop rules so cue indices line up with the stream's chapter
list, and degrades to the proportional split on any drift (script
edited after the render, stopped mid-book). Words are buttons —
click-to-seek, keyboard reachable — restyled as prose in index.css.
audiobook.lyrics translated in all 21 locales.
Co-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
* docs(changelog): synced-lyrics audiobook player entry (#1766)
Co-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
* style(audiobook): apply current formatter
* fix(audiobook): preserve synced render cues
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Palash Debnath <4178343+debpalash@users.noreply.github.com>
* feat(dub): project-level drag casting board
Adds an expandable Casting Board to the dub editor's CAST strip: speaker
rows (with the auto-clone chip when the extractor found a usable passage)
and draggable voice chips — Default, clone profiles, design presets.
Dropping a chip on a speaker writes the exact fields the CAST <select>
always has (profile_id + merge_parts/merge_parts_original attribution),
via a shared assignSpeakerProfile helper both views now call, so the
dropdowns stay in sync and job persistence is unchanged. Keyboard path:
focus a speaker row, pick from a listbox (arrows/Enter/Escape).
The pre-existing CAST dropdown strip moves verbatim into the new
CastingBoard.jsx (DubLeftColumn shrinks below 800 lines; the new file
holds the 300-line soft cap). Styles extend the .dub-cast-* cluster in
index.css. Six new i18n keys translated in all 21 locales.
Co-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
* docs: changelog + roadmap entries for the casting board (#1767)
Co-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
* fix(casting): validate and preserve speaker assignments
* fix(casting): recover cleared merged assignments
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Palash Debnath <4178343+debpalash@users.noreply.github.com>
Synchronize VoiceStudio release metadata, lockfiles, installers, container references, documentation, and the dated v0.5.2 changelog after all planned fixes landed.
Adds opt-in word-timed ASS karaoke captions while preserving the existing line-caption default.\n\nCo-authored-by: Matt Van Horn <mvanhorn@users.noreply.github.com>
An LLM agent pays for every byte it receives, and generate_speech returned
each WAV as base64 inline - a short clip already brushed per-result limits.
This adds two knobs in the OMNIVOICE_* family, the pattern the ElevenLabs
MCP settled on (OUTPUT_MODE + a BASE_PATH security boundary):
- OMNIVOICE_MCP_OUTPUT_MODE = resources (default, the original contract) |
files | both. In files mode generate_speech returns audio_url (the render
the backend already keeps, served at /audio/<id>.wav) and, when a base
path is set, output_path - the WAV written into that directory.
- OMNIVOICE_MCP_BASE_PATH: the one directory agents may read from and
receive files in. transcribe(audio_path=) and clone_voice(ref_audio_path=)
read only inside it (relative paths resolve against it, absolute ones must
lie within it, symlinks resolved before the check); with no base path,
path arguments are refused with a reason.
- OMNIVOICE_MCP_TIMEOUT_S (default 120): the tools' backend timeout, since a
CPU host serializes generations and an agent queued behind another render
outlasted the fixed budget with an empty-message ToolError.
Also: transcribe and clone_voice share one input helper (data-URI tolerance
now covers transcribe too), the upload filename carries the sniffed
extension, and the reply is built with json.dumps instead of hand-rolled
JSON. Tests cover the mode parsing, the boundary (escape and missing-base
refusals), both input lanes, all four reply shapes, and the timeout knob.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Replaces a scheduler-sensitive 200 ms assertion with a deterministic ordering check against the blocked-write release barrier. Repairs the red post-merge main run from #1751.
Refreshes README and linked docs with accurate installation, platform, privacy, API, and model-license guidance; documents local gigastt; and pins OpenAI-compatible ASR traffic to the configured secure origin. Closes#1736.