* fix(audiobook): chapter render no longer crashes on mixed 1-D/2-D audio chunks (#897)
Root cause: synthesize_chapter (backend/services/audiobook.py) built
inter-span pause silence as bare 1-D torch.zeros(n) while every real
engine's synth returns (1, samples) per the TTSBackend.generate contract
(OmniVoice's model.generate(...)[0] included) — so the chapter's final
hard concat in chunked_tts.concatenate_audio_chunks hit
torch.cat with mixed ranks and died with
'RuntimeError: Tensors must have same number of dimensions: got 1 and 2'.
Any chapter containing a [pause] span (Stories/audiobook longform)
crashed; existing tests missed it because their stub synth returned 1-D.
The crossfade branch had the same latent bug for mixed-rank chunks.
Fix, both layers:
- concatenate_audio_chunks now normalizes chunk shapes before any cat
(_normalize_chunk_shapes): lower-rank chunks gain leading singleton
dims to the highest rank present, then singleton channel dims
broadcast to the widest channel count (mono follows stereo). Covers
both the hard-cut and crossfade branches; homogeneous input passes
through untouched, so all-1-D / all-2-D callers keep their exact
output shapes. No future backend's output rank can re-break the join.
- synthesize_chapter materializes silence AFTER the loop, matching the
rendered audio's channel dims / dtype / device — the same pattern
generation.py's _render_with_pauses already uses for the single-shot
path — so the data is rank-consistent at the source too. A
silence-only chapter stays 1-D float32 as before.
Regression tests: mixed-rank hard-cut (both orders), mixed-rank
crossfade, mono->stereo broadcast, all-1-D/all-2-D shape stability, a
2-D-engine + [pause] chapter through synthesize_chapter (the exact #897
scenario), and a spy asserting the parts reaching the concat are
rank-homogeneous. All fail before the fix with the reported error.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): add the audiobook pause-span concat fix under [Unreleased] (#953)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
POST /audiobook/import ended with `plan.chapter_count`, but AudiobookPlan
only exposed `char_count` (a property) and emitted `chapter_count` from
`to_dict()` — so the attribute access raised AttributeError, surfacing as
"500 Internal Server Error: 'AudiobookPlan' object has no attribute
'chapter_count'". The parse itself succeeded, so this hit every import
format (.txt/.md/.epub/.pdf), not just PDF.
Add a `chapter_count` property mirroring `char_count`, and have `to_dict()`
derive its key from it so the attribute and serialized key can't drift.
No API/schema/data change.
Tests: unit property test + a direct-handler /audiobook/import regression
(pdf/md/txt) that fails-before with the AttributeError.
Fixes#543
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Inline delivery hints within a narration line, wired into BOTH front doors so
Audiobook and Stories behave identically.
- services/ssml_lite.py (parallel-built, 18 tests): parse_ssml_lite splits a
line into {text, speed, spell, emphasis} segments — nesting (innermost wins),
unclosed-to-EOL, stray-close ignored, adjacent-merge; ReDoS-safe literal
alternation. + spell_out().
- _parse_spans (audiobook script path) now applies SSML-lite as the innermost
layer (precedence: [voice:] → [pause] → SSML); each segment becomes a Span
with its speed (threaded to the renderer) and spelled-out text for [spell].
Trailing pause attaches to the run's last segment.
- frontend/src/utils/ssmlLite.js: client port (kept in sync with the .py) +
storyToSpans applies it per chunk — inline speed OVERRIDES the per-line slider,
falls back to it otherwise.
Tests: parse_ssml_lite (18 py + 10 js), script-level prosody parse, Stories
SSML compile (override + spell). 70 backend + 345 frontend green; build clean.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Lets a render correct hard-to-say words (e.g. {"GIF":"jiff","Dr":"Doctor"}).
Backend wiring; the editor UI folds into the full-width Audiobook redesign.
- services/pronunciation.py (parallel-built, 19 tests): apply_lexicon —
whole-word, case-insensitive, longest-first, word-boundary, single ReDoS-safe
re.sub pass; + normalize/load/save_lexicon (JSON).
- synthesize_chapter gains a `lexicon` kwarg, applied to each span's text before
chunk splitting (None/empty = no-op → backward compatible).
- _render_chapter_cached folds the normalized lexicon into the chapter cache key
(a lexicon edit re-renders); threaded through _render_longform_sse + the
/audiobook, /audiobook/preview, /longform/render request models.
Tests: synthesize_chapter respells via lexicon; pronunciation module (19);
75 related backend tests green.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
PR 5 moved Stories' full export to /longform/render but dropped per-line
**speed** — the old client export sent each line's speed to /generate; the
converged path silently ignored it. This restores it end-to-end.
- Span gains an optional `speed`; synthesize_chapter passes it to the injected
synth (signature now `synth(text, voice_id, speed)`); both engine paths
(OmniVoice model + generic TTSBackend) forward it to generate(speed=…).
- chapter_cache_key now includes speed (a speed change re-renders; tuples accept
an optional 4th element so existing 3-tuple callers/tests still work).
- LongformSpan + /longform/render carry speed; storyToSpans emits each line's
speed onto its spans.
Emotion note: per-line tone is already model-native via inline tags
([laughter] etc.) inserted into the text, so no separate emotion→instruct
plumbing is needed — the dead `emotion` store field stays unused/superseded.
Tests: storyToSpans speed passthrough (8); cache-key speed sensitivity; synth
stubs updated for the 3-arg signature. 65 backend + 334 frontend green; build clean.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Completes the audiobook backend: POST /audiobook renders each chapter through
the active TTS engine (synthesize_chapter + chunked_tts), writes per-chapter
WAVs, then muxes a chapterized m4b (FFMETADATA1 chapters via build_m4b_cmd +
concat demuxer). Progress streams as SSE (started/chapter/assembling/done/
error), recorded to job_store. ffmpeg-gated — emits an error event and stops
when ffmpeg is absent (m4b is the only output).
- services/audiobook.build_concat_list: pure ffmpeg concat-list builder with
proper single-quote escaping (no arg injection). Unit-tested.
- router: voice resolution (compact form of generation.py's locked/design/
clone cases) cached per id; OmniVoice native model path + generic TTSBackend
path; chapter synthesis runs on the GPU pool, ffmpeg via run_ffmpeg.
Reuses the tested building blocks from #402 (parser, synthesize_chapter,
FFMETADATA + m4b argv builders) — the new router glue is thin and
import-checked by CI. Deferred: epub/pdf ingest, ACX loudnorm mastering,
crash-resume, UI. 15 audiobook tests (added concat-list); docs §R3 updated.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(audiobook): chapterized audiobook core + plan preview (Wave 5)
First cut of the long-form vertical (parity §R3). Engine-agnostic core in
services/audiobook.py:
- parse_audiobook_script: pure parser. Markdown '# H1' headings → chapters;
inline [voice:NAME] switches the narrator; [pause …] is delegated to the
shared omnivoice.utils.text.parse_pause_markers so audiobooks and single-shot
synthesis keep one pause dialect. Returns a chapter/span plan.
- synthesize_chapter: orchestration via an injected synth(text, voice) callable
(reuses chunked_tts split + crossfade, stitches inter-span silence) — so it's
unit-testable with a stub backend, no model/GPU.
- build_chapter_ffmetadata + build_m4b_cmd: pure FFMETADATA1 [CHAPTER] builder
and faststart-m4b concat-demux argv (bitrate-validated, no injection).
POST /audiobook/plan returns the parsed plan (no TTS/ffmpeg, no side effects).
Deferred (follow-ups): the streaming synth job + chapterized-m4b run, epub/pdf
ingest (new dep), ACX loudnorm mastering, crash-resume, UI.
14 tests: parser (chapters/voice/pause/intro/empties/to_dict), FFMETADATA
offsets+escaping, m4b argv + bitrate guard, and stub-synth orchestration
(span+silence stitching, voice threading). docs §R3 status updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): linear-time regexes (CodeQL ReDoS)
CodeQL flagged polynomial backtracking on user-provided input in three
regexes reachable from the new POST /audiobook/plan endpoint:
- _VOICE_RE: \s*(...)\s* → single [^\]]* class, stripped in code.
- _HEADING_RE: trailing [ \t]* removed; title captured greedily + stripped.
- _PAUSE_RE (omnivoice/utils/text.py): the numeric spec is now an atomic
group (?>…) so its leading \s+ can't backtrack against the trailing \s*.
Behavior-preserving (Python >=3.11 already required); 14 pause tests + 14
audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): require non-space heading title start (CodeQL ReDoS)
The previous _HEADING_RE '[ \t]+(.+)' still let the leading whitespace class
and the title '.+' both match the same tab run (overlap → polynomial). Anchor
the title capture with \S so the two can't overlap. 14 audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): exclude '[' from voice-tag content (CodeQL ReDoS)
[^\]]* still matched '[', so a run of nested [voice: prefixes produced
overlapping finditer match attempts → O(n^2). Excluding both brackets
([^\]\[]) makes matches non-overlapping and linear. A voice name never
contains a bracket. 14 audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>