Commit Graph
5 Commits
Author SHA1 Message Date
Palash Debnath 85d8df2707 test(translator): stop patching time.sleep for the whole process (#1390)
test_chat_non_429_does_not_retry failed on an unrelated PR with assert not [30.0, 30.0, 30.0, ...] even though _chat had never slept. translator.py does a plain import time, so tr.time is the stdlib module and patching it replaces time.sleep process-wide: subprocess_backend's sidecar idle reaper, which loops on time.sleep(30.0), wrote into the assertion's list and busy-looped while the patch was held. Any test that spawns a sidecar armed it, so this was a latent flake for the whole suite.

Sleeps are now recorded per thread — the test's own waits are captured, everyone else's really sleep — through one helper every sleep-patching test in the file uses. Proven fails-before/passes-after with a background thread actively calling time.sleep.
2026-08-06 14:29:14 +05:30
46141c8e5e fix(dub): a rate-limited polish pass no longer skips fitting, fails the UI, or ignores Retry-After (#1135)
* fix(dub): a rate-limited polish pass no longer skips fitting, fails the UI, or ignores Retry-After

Observed live (owner's Bengali dub, 4 segments): every cinematic reflect call
429'd against a free-tier OpenRouter model and the UI declared "4/4 segment(s)
failed" over a translate that succeeded. Root-causing that surfaced a class,
not a message bug:

The cinematic reflect/adapt chain is OPTIONAL polish — on any failure the
segment keeps its literal translation and is fully usable. But every such
degradation (no-llm, reflect/adapt errors, adapt-diverged, wrong-script,
cinematic-budget) was reported under the same "error" key as real translation
failures. Three consumers took that at face value:

  1. useDubWorkflow counted the rows as failed -> the red N/N toast;
  2. _stamp_predicted_rate_ratio and _stamp_duration_plan skipped them ->
     no rate badges, no fits/tight/impossible verdicts;
  3. _apply_fit_pass and the condense pass skipped them -> overlong lines went
     to synthesis unfitted and came out audibly time-compressed at mix. This
     is a direct contributor to "later segments got worse" in rate-limited
     Cinematic dubs.

Split the vocabulary: "error" now means the row has no usable text (base
translation failed); optional-pass fallbacks ride a separate "degraded" key.
Downstream filters keep gating on "error" only, so degraded rows flow through
every fitting pass. The UI shows an amber "translated, polish skipped
(<reason>)" toast and a mild row tooltip instead of a red failure, and editing
a row clears the stale annotation.

And the retry that makes most of this moot: _chat now honors a 429's
Retry-After once (capped at 30s, jittered so the 6-wide segment fan-out does
not re-stampede the same window). OpenRouter's free pool says "Retry-After: 2"
- giving up instantly turned a two-second wait into a whole failed pass.

Tests: producer contract (every cinematic fallback returns degraded, never
error - 5 updated + retained), consumer contract (degraded rows still get
rate-ratio prediction and duration plans; error rows stay excluded), and the
retry (honors small Retry-After with jitter, caps absurd ones, one retry only,
non-429s never retry). Full suite: 2981 backend + 1236 frontend.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(changelog): correct PR ref to #1135

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(dub): review round — localize the degraded strings, un-suppress the mixed toast, clear stale annotations on edit

Three review findings, all valid:

- Localization parity (Greptile): the two new user-facing keys existed only in
  en.json. Every other key in these namespaces is translated in all 21
  locales, so the fallback-to-English behavior would have been a regression of
  the repo's parity convention. Both keys now translated in all 20 non-en
  locales, inserted beside their siblings.
- Mixed responses suppressed the degraded story (Greptile): when a translate
  returned both real failures and degraded rows, only the red failure toast
  fired. The degraded warning now fires alongside it — real failures don't
  erase what happened to the rows that succeeded plainly.
- Ordinary edits kept stale annotations (CodeRabbit): the restore path cleared
  translate_error/translate_degraded but a normal text edit didn't, so a row
  kept wearing "polish pass skipped" over words the user had just written.
  Editing the text now clears both annotations.

Frontend suite: 1236 passed; i18n probe green across all 21 locales.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 16:13:41 +05:30
437a995a0f fix(translate): Cinematic/Autofit can no longer invent dialogue — divergence guard + pinned temperature (#950)
Root cause (v0.3.9 field report): the refine paths had no-op output guards.
Cinematic's ADAPT step only checked _looks_like_target_script, which returns
True unconditionally for every Latin-script target (no _SCRIPT_RANGES entry)
— so any non-empty LLM reply (hallucinated dialogue, refusals, commentary,
or the REFLECT critique itself) shipped as the dub line. Autofit's
adjust_for_slot accepted ANY non-empty reply, and its best-candidate tracker
(closest rate_ratio to 1.0) actively selected the most-padded output, while
_EXPAND_PROMPT invited invention with no ceiling. Both call paths also ran
at the provider-default temperature 1.0, unlike the working Fast path which
pins 0.2.

The fix, class-level:
- Shared divergence guard translator.refine_output_ok (length window
  0.4–2.5x, env-tunable via OMNIVOICE_REFINE_RATIO_MIN/MAX, with an
  absolute cap for short references; target-script check; critique-echo
  detection). Rejected ADAPT output degrades to the literal with
  error="adapt-diverged" (wrong-script keeps its adapt-wrong-script:<lang>
  marker), riding the existing degradation machinery unchanged.
- Autofit validates every reply against the ORIGINAL input text (divergence
  compounds across attempts otherwise); rejected candidates are discarded
  (attempt burned, graceful degradation to the input preserved) with
  error="fit-diverged"; lines under 15% of their slot skip LLM expansion
  entirely (fit-skip-short) — they could only "fill" the slot with
  fabricated dialogue.
- temperature=0.2 pinned on the cinematic (_chat) and fit (llm.chat) calls;
  chat/chat_messages gained an optional temperature param that is only sent
  when set, so refinement/director/glossary callers keep provider defaults.
- Prompts hardened: ADAPT forbids introducing facts/names/dialogue not in
  the source line; EXPAND forbids inventing information and more than
  doubling the line.
- speech_rate strict-mode docstring made honest: strict changes only the
  upper tolerance bound; expansion still runs (now guard-bounded).

Fail-before/pass-after regression tests for the reported bugs (10x runaway
ADAPT on an es target, critique echo, hallucinated slot-fill expansion,
refusal replies, tiny-line expansion skip, pinned temperature) plus the
previously-untested wrong-script fallback and legit-output acceptance.

Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 22:28:37 +05:30
Palash Debnathandmergetest 86c701bff9 fix(translate): bound the whole cinematic/autofit pass so a slow LLM can't hang "Translating…" (#868)
The per-segment call already has a 45s timeout and concurrency is capped, but a
slow or rate-limited provider on a large dub (hundreds of segments) can still keep
the "Translating…" spinner spinning for minutes as segments queue through the
bounded pool. There was no ceiling on the *whole* pass.

Add an overall wall-clock budget, OMNIVOICE_CINEMATIC_BUDGET_S (default 180s,
<=0 disables). Segments that finish in time keep their cinematic refine; any still
in-flight when the budget hits is cancelled and degrades to its literal (Fast)
translation with error="cinematic-budget", so the translate ALWAYS returns instead
of hanging. Order and length of the result are preserved. Abandoned executor
threads follow the same fire-and-forget pattern as the GPU-pool wedge guard (#730).

Regression tests: a 3s-per-segment refine under a 0.3s budget returns in <2s with
literal fallbacks; budget<=0 runs every segment to completion.

Co-authored-by: mergetest <test@local>
2026-07-02 00:39:36 +05:30
debpalash 52d68d05dc refactor: update backend architecture, expand frontend state management, and synchronize voice-pro research modules. 2026-04-21 18:32:25 +05:30