- validate managed binaries before exec; 0-byte GGUF placeholders fail
actionably instead of "Exec format error" (#1172)
- KittenTTS: tokenizer-measured chunking to the ONNX 512-token cap;
clear 400 for unspeakable input (#1173)
- clean SIGTERM during weight load: shutdown-aware loader, benign
cancelled-load classification, lifespan hardening, scoped log
silencers (transformers load + alembic fileConfig) (#1174)
- broken ASR deep-imports (lightning_fabric) mark the engine
unavailable with a repair hint and fall through (#1185)
- uv cache + managed Python follow the chosen install drive on
Windows (cherry-picked cross-drive class fix + spaces/D: tests) (#1186)
- adaptive silence-removal ladder for quiet clone references; localized
actionable error for truly silent clips, all 21 locales (#1188)
- CHANGELOG: consolidated Unreleased into the quiet one-liner style
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(design): heal validator-rejecting instruct on design voices (#594/#571/#596)
A designed voice could persist an `instruct` the engine validator rejects —
either the literal "[object Object]" from a pre-fix build (#550) or freeform
prose typed into the style field — so every Generate/Dub that used the voice
failed with `Unsupported instruct items found in …` (400/500, and "Can't reach
the local backend" when it tore down mid-render). Migration 0006 only *blanked*
"[object Object]", which silently discarded the design — an Indonesian female
voice then rendered male (#594).
Fix the whole class by healing at every seam and rebuilding from the
authoritative source (the design's saved `vd_states` category picks):
- omnivoice/utils/voice_design.py: add sanitize_instruct / instruct_from_vd_states
/ heal_design_instruct — forgiving (never raise), drop poison/prose to valid
tags, and rebuild tags from vd_states when the stored value is unusable.
- profiles.py: sanitize + rebuild at save (POST) and sanitize at edit (PUT), so
no poisoned instruct can ever be persisted again.
- generation.py + dub_generate.py: heal whenever a profile drives synthesis, so
legacy poisoned rows resolve to valid tags instead of 400-ing.
- migration 0007: heal existing profiles in place (recovers gender/age/pitch
from vd_states), self-contained (frozen vocab snapshot) so it never drags
torch into startup; supersedes 0006's blanking. Backward-compatible.
Tests: unit coverage for the healer, a migration test driving 0006->0007 on the
real schema, a parity guard so the frozen snapshot can't drift, and two API
guards. Corrected one existing test that had encoded the #594 behaviour.
Resolves#571, #594, #596; removes a major driver of the "Can't reach backend"
reports.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(cjk): allowlist migration 0007's frozen dialect-tag snapshot (#564)
The 0007 instruct-heal migration carries a frozen copy of the design-tag
whitelist (incl. Chinese dialect tags) so it stays self-contained; add it to
the hardcoded-CJK allowlist like omnivoice/utils/voice_design.py.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Non-English voices drifted to English/wrong-language because the request's or
profile's language wasn't reaching the model:
- #533: generate_speech() read instruct/ref_text/seed from a resolved profile row
but never row['language'] (and collapsed Auto→None), so a German archetype
previewed in German yet generated in English on the user's own call (and via
Docker/API). Fall back to the profile's stored language when the request didn't
pin one; an explicit non-Auto request language still wins. (Frontend already
sets the dropdown on profile-select; this is the authoritative backend fix.)
- #505 (B2): the audiobook/longform synth hardcoded language=None, so the engine
re-autodetected per chunk and a non-English clone flipped language mid-render.
Add _resolve_default_language (request → profile → autodetect) and thread the
resolved language through _build_synth/_prepare_synth/_render_longform_sse, the
three longform request models, the preview path, and the resume manifest.
Genuine Auto/unset behavior is unchanged.
- #502 (partial): the duration estimator weights combining marks (U+0300–036F)
at 0.0, so NFD/decomposed text under-allocated frames → rushed audio. NFC-
normalize text at the estimator entry — fixes the whole diacritic-script class
(no-op for precomposed text). (The residual "distorted" core still needs the
reporter's sample; tracked separately.)
Tests (fail-before/pass-after): profile language reaches the engine (German→de;
explicit/Auto override semantics); longform synth gets the resolved language
(→ja), not None; NFD vs NFC duration parity (Korean Hangul diverges ~3x pre-fix).
Full suite: tests/ 1740 passed, backend/tests/ 114 passed.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(audiobook): chapterized audiobook core + plan preview (Wave 5)
First cut of the long-form vertical (parity §R3). Engine-agnostic core in
services/audiobook.py:
- parse_audiobook_script: pure parser. Markdown '# H1' headings → chapters;
inline [voice:NAME] switches the narrator; [pause …] is delegated to the
shared omnivoice.utils.text.parse_pause_markers so audiobooks and single-shot
synthesis keep one pause dialect. Returns a chapter/span plan.
- synthesize_chapter: orchestration via an injected synth(text, voice) callable
(reuses chunked_tts split + crossfade, stitches inter-span silence) — so it's
unit-testable with a stub backend, no model/GPU.
- build_chapter_ffmetadata + build_m4b_cmd: pure FFMETADATA1 [CHAPTER] builder
and faststart-m4b concat-demux argv (bitrate-validated, no injection).
POST /audiobook/plan returns the parsed plan (no TTS/ffmpeg, no side effects).
Deferred (follow-ups): the streaming synth job + chapterized-m4b run, epub/pdf
ingest (new dep), ACX loudnorm mastering, crash-resume, UI.
14 tests: parser (chapters/voice/pause/intro/empties/to_dict), FFMETADATA
offsets+escaping, m4b argv + bitrate guard, and stub-synth orchestration
(span+silence stitching, voice threading). docs §R3 status updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): linear-time regexes (CodeQL ReDoS)
CodeQL flagged polynomial backtracking on user-provided input in three
regexes reachable from the new POST /audiobook/plan endpoint:
- _VOICE_RE: \s*(...)\s* → single [^\]]* class, stripped in code.
- _HEADING_RE: trailing [ \t]* removed; title captured greedily + stripped.
- _PAUSE_RE (omnivoice/utils/text.py): the numeric spec is now an atomic
group (?>…) so its leading \s+ can't backtrack against the trailing \s*.
Behavior-preserving (Python >=3.11 already required); 14 pause tests + 14
audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): require non-space heading title start (CodeQL ReDoS)
The previous _HEADING_RE '[ \t]+(.+)' still let the leading whitespace class
and the title '.+' both match the same tab run (overlap → polynomial). Anchor
the title capture with \S so the two can't overlap. 14 audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): exclude '[' from voice-tag content (CodeQL ReDoS)
[^\]]* still matched '[', so a run of nested [voice: prefixes produced
overlapping finditer match attempts → O(n^2). Excluding both brackets
([^\]\[]) makes matches non-overlapping and linear. A voice name never
contains a bracket. 14 audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Lets users insert pauses in the transcript: `[pause]` (350ms default),
`[pause 500ms]`, `[pause 1s]`, `[pause 1.5s]`. Requester confirmed the
`[pause Nms]` syntax (fits the existing marker style).
Implementation is fully opt-in and model-free:
- `omnivoice/utils/text.parse_pause_markers()` splits the text into
`(span, pause_ms_after)` tuples (case-insensitive; bare number = ms; `s`
suffix = seconds; adjacent markers sum; clamped to 10s). Text with no marker
returns unchanged, so existing behavior is untouched.
- `_run_inference` synthesizes each span as today and stitches a `torch.zeros`
silence buffer between them at the `[pause]` points (matching channel
dims/dtype/device); DSP/mastering then runs once over the combined audio.
An explicit overall `duration` isn't split across spans (left to the model
per span).
Tests (no TTS model loaded): tests/test_pause_markers.py covers the parser
(ms/s/default/clamp/leading/trailing/adjacent/round-trip) and the silence
stitching with a fake gen fn (lengths + zeroed regions). Full pause + CJK guard
+ router smoke suites pass (39).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>