* feat(audiobook): chapterized audiobook core + plan preview (Wave 5)
First cut of the long-form vertical (parity §R3). Engine-agnostic core in
services/audiobook.py:
- parse_audiobook_script: pure parser. Markdown '# H1' headings → chapters;
inline [voice:NAME] switches the narrator; [pause …] is delegated to the
shared omnivoice.utils.text.parse_pause_markers so audiobooks and single-shot
synthesis keep one pause dialect. Returns a chapter/span plan.
- synthesize_chapter: orchestration via an injected synth(text, voice) callable
(reuses chunked_tts split + crossfade, stitches inter-span silence) — so it's
unit-testable with a stub backend, no model/GPU.
- build_chapter_ffmetadata + build_m4b_cmd: pure FFMETADATA1 [CHAPTER] builder
and faststart-m4b concat-demux argv (bitrate-validated, no injection).
POST /audiobook/plan returns the parsed plan (no TTS/ffmpeg, no side effects).
Deferred (follow-ups): the streaming synth job + chapterized-m4b run, epub/pdf
ingest (new dep), ACX loudnorm mastering, crash-resume, UI.
14 tests: parser (chapters/voice/pause/intro/empties/to_dict), FFMETADATA
offsets+escaping, m4b argv + bitrate guard, and stub-synth orchestration
(span+silence stitching, voice threading). docs §R3 status updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): linear-time regexes (CodeQL ReDoS)
CodeQL flagged polynomial backtracking on user-provided input in three
regexes reachable from the new POST /audiobook/plan endpoint:
- _VOICE_RE: \s*(...)\s* → single [^\]]* class, stripped in code.
- _HEADING_RE: trailing [ \t]* removed; title captured greedily + stripped.
- _PAUSE_RE (omnivoice/utils/text.py): the numeric spec is now an atomic
group (?>…) so its leading \s+ can't backtrack against the trailing \s*.
Behavior-preserving (Python >=3.11 already required); 14 pause tests + 14
audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): require non-space heading title start (CodeQL ReDoS)
The previous _HEADING_RE '[ \t]+(.+)' still let the leading whitespace class
and the title '.+' both match the same tab run (overlap → polynomial). Anchor
the title capture with \S so the two can't overlap. 14 audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(audiobook): exclude '[' from voice-tag content (CodeQL ReDoS)
[^\]]* still matched '[', so a run of nested [voice: prefixes produced
overlapping finditer match attempts → O(n^2). Excluding both brackets
([^\]\[]) makes matches non-overlapping and linear. A voice name never
contains a bracket. 14 audiobook tests green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>