test_chat_non_429_does_not_retry failed on an unrelated PR with assert not [30.0, 30.0, 30.0, ...] even though _chat had never slept. translator.py does a plain import time, so tr.time is the stdlib module and patching it replaces time.sleep process-wide: subprocess_backend's sidecar idle reaper, which loops on time.sleep(30.0), wrote into the assertion's list and busy-looped while the patch was held. Any test that spawns a sidecar armed it, so this was a latent flake for the whole suite.
Sleeps are now recorded per thread — the test's own waits are captured, everyone else's really sleep — through one helper every sleep-patching test in the file uses. Proven fails-before/passes-after with a background thread actively calling time.sleep.
* fix(dub): a rate-limited polish pass no longer skips fitting, fails the UI, or ignores Retry-After
Observed live (owner's Bengali dub, 4 segments): every cinematic reflect call
429'd against a free-tier OpenRouter model and the UI declared "4/4 segment(s)
failed" over a translate that succeeded. Root-causing that surfaced a class,
not a message bug:
The cinematic reflect/adapt chain is OPTIONAL polish — on any failure the
segment keeps its literal translation and is fully usable. But every such
degradation (no-llm, reflect/adapt errors, adapt-diverged, wrong-script,
cinematic-budget) was reported under the same "error" key as real translation
failures. Three consumers took that at face value:
1. useDubWorkflow counted the rows as failed -> the red N/N toast;
2. _stamp_predicted_rate_ratio and _stamp_duration_plan skipped them ->
no rate badges, no fits/tight/impossible verdicts;
3. _apply_fit_pass and the condense pass skipped them -> overlong lines went
to synthesis unfitted and came out audibly time-compressed at mix. This
is a direct contributor to "later segments got worse" in rate-limited
Cinematic dubs.
Split the vocabulary: "error" now means the row has no usable text (base
translation failed); optional-pass fallbacks ride a separate "degraded" key.
Downstream filters keep gating on "error" only, so degraded rows flow through
every fitting pass. The UI shows an amber "translated, polish skipped
(<reason>)" toast and a mild row tooltip instead of a red failure, and editing
a row clears the stale annotation.
And the retry that makes most of this moot: _chat now honors a 429's
Retry-After once (capped at 30s, jittered so the 6-wide segment fan-out does
not re-stampede the same window). OpenRouter's free pool says "Retry-After: 2"
- giving up instantly turned a two-second wait into a whole failed pass.
Tests: producer contract (every cinematic fallback returns degraded, never
error - 5 updated + retained), consumer contract (degraded rows still get
rate-ratio prediction and duration plans; error rows stay excluded), and the
retry (honors small Retry-After with jitter, caps absurd ones, one retry only,
non-429s never retry). Full suite: 2981 backend + 1236 frontend.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(changelog): correct PR ref to #1135
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(dub): review round — localize the degraded strings, un-suppress the mixed toast, clear stale annotations on edit
Three review findings, all valid:
- Localization parity (Greptile): the two new user-facing keys existed only in
en.json. Every other key in these namespaces is translated in all 21
locales, so the fallback-to-English behavior would have been a regression of
the repo's parity convention. Both keys now translated in all 20 non-en
locales, inserted beside their siblings.
- Mixed responses suppressed the degraded story (Greptile): when a translate
returned both real failures and degraded rows, only the red failure toast
fired. The degraded warning now fires alongside it — real failures don't
erase what happened to the rows that succeeded plainly.
- Ordinary edits kept stale annotations (CodeRabbit): the restore path cleared
translate_error/translate_degraded but a normal text edit didn't, so a row
kept wearing "polish pass skipped" over words the user had just written.
Editing the text now clears both annotations.
Frontend suite: 1236 passed; i18n probe green across all 21 locales.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Root cause (v0.3.9 field report): the refine paths had no-op output guards.
Cinematic's ADAPT step only checked _looks_like_target_script, which returns
True unconditionally for every Latin-script target (no _SCRIPT_RANGES entry)
— so any non-empty LLM reply (hallucinated dialogue, refusals, commentary,
or the REFLECT critique itself) shipped as the dub line. Autofit's
adjust_for_slot accepted ANY non-empty reply, and its best-candidate tracker
(closest rate_ratio to 1.0) actively selected the most-padded output, while
_EXPAND_PROMPT invited invention with no ceiling. Both call paths also ran
at the provider-default temperature 1.0, unlike the working Fast path which
pins 0.2.
The fix, class-level:
- Shared divergence guard translator.refine_output_ok (length window
0.4–2.5x, env-tunable via OMNIVOICE_REFINE_RATIO_MIN/MAX, with an
absolute cap for short references; target-script check; critique-echo
detection). Rejected ADAPT output degrades to the literal with
error="adapt-diverged" (wrong-script keeps its adapt-wrong-script:<lang>
marker), riding the existing degradation machinery unchanged.
- Autofit validates every reply against the ORIGINAL input text (divergence
compounds across attempts otherwise); rejected candidates are discarded
(attempt burned, graceful degradation to the input preserved) with
error="fit-diverged"; lines under 15% of their slot skip LLM expansion
entirely (fit-skip-short) — they could only "fill" the slot with
fabricated dialogue.
- temperature=0.2 pinned on the cinematic (_chat) and fit (llm.chat) calls;
chat/chat_messages gained an optional temperature param that is only sent
when set, so refinement/director/glossary callers keep provider defaults.
- Prompts hardened: ADAPT forbids introducing facts/names/dialogue not in
the source line; EXPAND forbids inventing information and more than
doubling the line.
- speech_rate strict-mode docstring made honest: strict changes only the
upper tolerance bound; expansion still runs (now guard-bounded).
Fail-before/pass-after regression tests for the reported bugs (10x runaway
ADAPT on an es target, critique echo, hallucinated slot-fill expansion,
refusal replies, tiny-line expansion skip, pinned temperature) plus the
previously-untested wrong-script fallback and legit-output acceptance.
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The per-segment call already has a 45s timeout and concurrency is capped, but a
slow or rate-limited provider on a large dub (hundreds of segments) can still keep
the "Translating…" spinner spinning for minutes as segments queue through the
bounded pool. There was no ceiling on the *whole* pass.
Add an overall wall-clock budget, OMNIVOICE_CINEMATIC_BUDGET_S (default 180s,
<=0 disables). Segments that finish in time keep their cinematic refine; any still
in-flight when the budget hits is cancelled and degrades to its literal (Fast)
translation with error="cinematic-budget", so the translate ALWAYS returns instead
of hanging. Order and length of the result are preserved. Abandoned executor
threads follow the same fire-and-forget pattern as the GPU-pool wedge guard (#730).
Regression tests: a 3s-per-segment refine under a 0.3s budget returns in <2s with
literal fallbacks; budget<=0 runs every segment to completion.
Co-authored-by: mergetest <test@local>