Root cause (v0.3.9 field report): the refine paths had no-op output guards.
Cinematic's ADAPT step only checked _looks_like_target_script, which returns
True unconditionally for every Latin-script target (no _SCRIPT_RANGES entry)
— so any non-empty LLM reply (hallucinated dialogue, refusals, commentary,
or the REFLECT critique itself) shipped as the dub line. Autofit's
adjust_for_slot accepted ANY non-empty reply, and its best-candidate tracker
(closest rate_ratio to 1.0) actively selected the most-padded output, while
_EXPAND_PROMPT invited invention with no ceiling. Both call paths also ran
at the provider-default temperature 1.0, unlike the working Fast path which
pins 0.2.
The fix, class-level:
- Shared divergence guard translator.refine_output_ok (length window
0.4–2.5x, env-tunable via OMNIVOICE_REFINE_RATIO_MIN/MAX, with an
absolute cap for short references; target-script check; critique-echo
detection). Rejected ADAPT output degrades to the literal with
error="adapt-diverged" (wrong-script keeps its adapt-wrong-script:<lang>
marker), riding the existing degradation machinery unchanged.
- Autofit validates every reply against the ORIGINAL input text (divergence
compounds across attempts otherwise); rejected candidates are discarded
(attempt burned, graceful degradation to the input preserved) with
error="fit-diverged"; lines under 15% of their slot skip LLM expansion
entirely (fit-skip-short) — they could only "fill" the slot with
fabricated dialogue.
- temperature=0.2 pinned on the cinematic (_chat) and fit (llm.chat) calls;
chat/chat_messages gained an optional temperature param that is only sent
when set, so refinement/director/glossary callers keep provider defaults.
- Prompts hardened: ADAPT forbids introducing facts/names/dialogue not in
the source line; EXPAND forbids inventing information and more than
doubling the line.
- speech_rate strict-mode docstring made honest: strict changes only the
upper tolerance bound; expansion still runs (now guard-bounded).
Fail-before/pass-after regression tests for the reported bugs (10x runaway
ADAPT on an es target, critique echo, hallucinated slot-fill expansion,
refusal replies, tiny-line expansion skip, pinned temperature) plus the
previously-untested wrong-script fallback and legit-output acceptance.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Items 1 and 2 from the improvement list:
1. Synchronization — Whisper-family ASR stretches segment starts back
over leading non-speech (intro music, silence), so the dub starts at
0:00 while the speaker starts at 0:02-0:03. New onset_align service
snaps each segment start forward to the first audible vocal onset
(adaptive RMS threshold over the Demucs-isolated vocals when
available). Forward-only and conservative: never moves a start
earlier, ignores sub-100ms shifts, preserves minimum duration,
leaves silent-window segments untouched. Pure NumPy — identical
across platforms.
2. Accent/vocabulary by country — a Dialect picker in the Dub panel
(BCP-47 codes per target language) injects a regional instruction
into LLM translation prompts (OpenAI/Ollama engines and the
Cinematic refine pass): Argentina yields 'Vos sos muy listo', not
'Tú eres muy listo'. Non-LLM engines show a clear hint that the
dialect needs an LLM. New i18n keys translated in all 21 locales.
Item 3 (segment rectangles: move/crop/stretch on the timeline) is a
larger editor feature and stays open on #280.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: mergetest <test@local>