POST /v1/audio/speech returned synthetic audio without the AudioSeal
provenance watermark while /generate marked the same text — and the
audit behind #1169 found the same class of gap in five more producers.
EU AI Act Art. 50(2) (applicable 2026-08-02, expressly carved out of
the open-source exemption in Art. 2(12)) makes machine-readable marking
of synthetic audio a provider obligation, so per-door coverage gaps are
compliance bugs.
Root cause: coverage grew call-site-by-call-site (three separate
embed_watermark calls) with nothing forcing a new producer to opt in.
The fix, per the whole-class rule:
* services/watermark.py grows `mark_synthetic(wav, sr, *, context,
force=False)` — THE named chokepoint (delegates to embed_watermark:
pref-gated, no-op without AudioSeal, never raises / degrades to
unmarked) — plus `will_mark()` for cache-key derivation.
* Gaps closed at the tensor stage, before any encoding:
- openai_compat `_run_tts` (the reported gap; all response_formats)
- tts_stream /ws/tts (per-sentence, before PCM16 framing)
- generation stream=true preview chunks (marked streamed copy; the
saved take keeps its single whole-take mark in finalize)
- batch dub pipeline (assembled track, before WAV write / aac mux)
- longform chapter render — covers /audiobook, /longform/render
(Stories), /audiobook/preview and /audiobook/resume; the chapter
cache key now carries a watermark tag so stale unmarked cache
entries can never be served for a marked-on render
- dub preview-segment (docstring had declared it exempt)
- archetype render (served preview + materialized profile reference)
* Already-covered paths (generate finalize, dub segments, persona
bundles) migrated onto the same chokepoint.
* Documented non-producers/gaps instead of fake coverage:
/stories/encode is a pure transcoder of user uploads (must NOT mark);
the opt-in SoniTranslate sidecar synthesizes+muxes externally and is
a documented provenance gap at its route.
* Regression tests: tests/test_synthetic_audio_watermark_1169.py runs
detect_watermark() on the actual response audio of every producing
route (real watermark service, fake AudioSeal nets, fake engine;
verified fail-before on the pre-fix tree — 11 of 13 fail).
* Recurrence-proofing: tests/test_watermark_route_coverage.py
structurally asserts every synthesis call site references
mark_synthetic (justified allowlist), producers keep their call,
and embed_watermark is never called outside the chokepoint.
Settings toggle semantics are unchanged (the issue's two legal
questions stay flagged for a lawyer, deliberately unanswered here).
New pure planning layer (services/duration_planner.py) runs after translation,
before TTS: estimates each translated line's natural speech duration (self-
calibrating from the job's already-synthesized segments, static per-language
rates as cold-start fallback) and classifies it fits/tight/impossible against
slot + capped gap borrow, with thresholds derived from fit_planner's own caps
so "impossible" means "would be trimmed". Verdicts ride the /dub/translate
response and badge the segment table; an opt-in (default OFF) LLM pass attaches
one-click shorter-rewrite suggestions for impossible lines. Never blocks
generation — informs before GPU time is burned.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>