The compute-time error told users to raise a generation timeout that had no control anywhere in the app — the only knob was an environment variable, and on Windows the docs explicitly warn against the usual way of setting one. Both budgets are now editable in Settings under Performance & Device, persisted and applied on the next start.
Two defects found in review and fixed here rather than shipped: an explicit universal budget silently overrode a separately saved CPU budget, so the CPU row would have looked like it worked and done nothing; and a value already set in the environment shadowed the saved preference while the panel still reported success. A shadowed row now says so instead. Long-input warnings also fire on Apple Silicon, which gets the accelerated budget and was the device in one of the duplicate reports.
Fixes#1787. Closes the reports tracked in #1774 and #1778.
CodeRabbit on 28c7bace:
1. (Major) get_watermark_pool's double-checked pattern re-read the
global after an unlocked null-check, so shutdown_watermark_pool's
reset could land in between and the caller received None. The
executor is now captured and returned under _watermark_pool_lock.
2. (Minor) the idle-grace test overwrote _prefetched_unused after the
embed call, making the embed's clearing unobservable — a failing
embed would have passed unnoticed. It now asserts the flag directly,
and a guard diverts any leaked idle reaper (idle_worker resolves
release_idle_models per call) to a no-op for the test's duration.
The shutdown drain killed the module singleton with no replacement, so
any process that keeps running after a lifespan shutdown — the CI suite
does exactly this — dead-submitted on the next watermark op: "cannot
schedule new futures after shutdown" (CI red; independently confirmed
by Greptile P1, CodeRabbit Major, and the plugin code review at 95/100
confidence). shutdown_watermark_pool() now resets the singleton under
its build lock before draining, so the next get_watermark_pool() hands
out a live replacement. Regression test covers
drained-pool-refuses + replacement-accepts.
Same round, minor findings: the drain's except now logs with exc_info
instead of a bare pass (GHAS CodeQL empty-except); the watermark delay
knob rejects negative/non-finite overrides (CodeRabbit); conftest sets
OMNIVOICE_PRELOAD_WATERMARK=0 unconditionally so a stray export from
the runner shell cannot re-enable background warm-ups mid-suite
(CodeRabbit).
* feat(omnivoice): port upstream VoiceClonePrompt persistence + FlashInfer opt-in
Upstream k2-fsa teardown ports, verified with generated voice samples:
- VoiceClonePrompt.save()/.load() (upstream format v1, weights_only-safe)
on the vendored model, and a disk layer under the in-memory prompt LRU
(DATA_DIR/prompt_cache, keyed by ref path+mtime+ref_text+preprocess,
32 newest kept, OMNIVOICE_PROMPT_DISK_CACHE=0 opts out). First generation
of a session with a known voice skips the reference re-encode and any
auto-transcription pass — verified across two real processes (encodes=1
then encodes=0, same voice).
- omnivoice_flashinfer.py ported (packed CFG attention, fused kernels,
optional CUDA graphs), schedule adapted to our num_step+1 divergence.
Opt-in via OMNIVOICE_FLASHINFER=1|graph, CUDA-only, replaces
torch.compile for the session; missing package / apply failure / runtime
failure all degrade with a named reason (same #278 contract as compile:
classify → unapply → retry once, session latch). Measured 2.20x at
batch=1 on an RTX 4090 with byte-identical text and clean ASR round-trip.
- Docs: OmniVoice guide gains instruct+reference combination semantics
(consistent instruct stabilizes cloning, reference wins conflicts),
inline pronunciation control (pinyin / CMU), prompt persistence, and
corrects the 'no voice design' claim; performance.md documents both new
env knobs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: point changelog entries at the real PR number (#1565)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(pr): harden FlashInfer lifecycle + prompt-cache writes per review
Bot harvest round 1 (#1565): unapply on apply-failure (half-patched model
could crash the next render); pin eager-mode FlashInfer inference to one
thread too — the attention plan and packed position ids are per-generation
module state, so interleaved _gpu_pool workers would corrupt each other;
restore the CAPTURED pre-apply attention impl (could be flash_attention_2)
instead of assuming sdpa; unique tmp name per prompt-cache write; correct
the _forward_logits layout docstring; resolve VoiceClonePrompt at test
runtime; docs — Known limits keeps only the limitation, performance.md
states the VRAM cost and scopes the fallback claim to classified kernel
failures.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(pr): round-2 review — publish only a fully restored model, redact latch reason, tighten CPU-persistence test
Greptile: the runtime fallback now unapplies BEFORE swapping generate, so
a concurrent render keeps queuing behind the thread-affinity wrapper while
teardown mutates modules. CodeRabbit: FlashInfer failure reasons pass
through core.failure.sanitize before latching/logging (wheel paths embed
the user's home); the save-portability test now creates the tokens on CUDA
when available and asserts the persisted payload itself is CPU-resident.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(pr): fail-closed latch reason when the sanitizer itself breaks
CodeQL empty-except + CodeRabbit round 3: if core.failure.sanitize raises,
the raw reason (home paths, wheel paths) was latched anyway. Now only the
exception class survives with a fixed redaction note; two regression tests
(normal redaction + sanitizer failure).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* refactor(launchpad): quieter, borderless design refresh
The launchpad carried decoration from an earlier direction: icon chips,
corner-hung count badges, a permanently visible filled arrow, uppercase
mono card titles, and a dotted stipple divider — plus a frame that had
been invisible since the app-wide border tokens were zeroed.
Rework it around what the borderless direction actually implies:
- Feature tiles get a whisper-faint surface instead of a dead frame, and
read as three bands (bare glyph + count / title + arrow / description).
`--card-hue` is spent sparingly — the glyph at rest, the surface, count
and arrow only once raised. Titles move to sans sentence case; counts
are plain tabular numerals. Lift softened 4px -> 2px, coloured glow ->
neutral shadow, plus an explicit focus ring and a staggered entrance.
- Hero drops the boxed "646" pill and the filled A/B-Compare button for
quiet type, with a hairline standing in for the separation.
- Section labels trade the dotted stipple for a single fading hairline;
rows are transparent until hover and reveal "Open" on hover/focus (it
stays in the DOM, so AT and keyboard always reach it).
- Hero, tiles, recent files, callout and project lists now share one
1180px column — previously only the top half was capped, so lists ran
edge-to-edge on a wide display while the deck stayed centred.
Two bugs found and fixed while doing it:
- Buttons that had `border border-solid border-transparent` removed fell
back to the UA default border and rendered a visible 1px outline. They
now carry `border-0` explicitly.
- `.lp-animate` used `animation-fill-mode: both`, so after the entrance
it kept owning `transform` — and animation-origin declarations outrank
normal ones, which silently killed the card hover lift. Now `backwards`,
which still holds the from-state through the stagger delay.
Also drops CSS the page has not rendered since #904: the cursor-spotlight
layer, the breath ring, and the per-card waveform strip.
Verified with headless renders at 1600/1280/940 and the empty state.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(dictation): decode Wayland portal signals and show the capture pill
The GlobalShortcuts portal declares Activated/Deactivated as
(o session, s shortcut_id, t timestamp, a{sv} options). We decoded the
timestamp as u32, so zbus rejected every signal with
Signature mismatch: got `(osta{sv})`, expected `(osua{sv})`
and the press was dropped as an invalid signal. Registration succeeded
and the desktop even reported the bound chord back, so the hotkey looked
wired up while doing nothing at all — on every Wayland compositor, for
the whole life of the feature (#1490). Decode the 64-bit timestamp, and
keep the 32-bit spelling as a fallback so a non-conforming portal
degrades to working rather than to silence.
With presses arriving, the second half of the failure showed: nothing
had shown the widget window since it became a hidden recorder host, so a
capture ran with no pill on screen — and a mic or Accessibility failure
rendered into a window nobody could see. Add show_dictation_pill, which
bottom-centres the capsule on the monitor under the pointer and shows it
without taking focus (Windows keeps SW_SHOWNOACTIVATE so paste still
lands in the user's document), and call it from the widget for every
state but idle. Wayland denies clients their own placement, so the
compositor picks the spot there; the pill still appears.
dispatch_dictation_capture now logs whether a press was emitted or
queued — a press that reaches Rust and produces nothing was otherwise
indistinguishable from one the compositor never delivered.
Tests: portal signals decode at both timestamp widths (the 64-bit case
fails before this change with the exact production error); pill
placement centres, respects a second monitor's origin, and clamps rather
than going off-screen; the widget shows for a state needing the user,
stays hidden while idle, and never shows for a press that arrives while
dictation is disabled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: sync in-progress workspace changes
Uncommitted work already in the tree, checkpointed so the branch matches
the local machine:
- Remote GPU workers: join-from-the-app flow, one-time secrets, QR join
codes, a Compute control in the status bar, and the device-list
Workers panel (#1516)
- Model Catalogue workspace, with Settings pointing at it
- Settings sidebar search and keyboard navigation
- Demo assets for dubbing, dictation and voice design, plus the scripts
that render them
- Backend: validation-error handling, ASR request-path degradation, and
the accompanying tests
- CHANGELOG entries for the above and for the Wayland dictation fix
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(tests): follow Engines to the Model Catalogue, and green the sweep
- test_supertonic3 asserted the license gate points at "Settings" while
the engine now names Model Catalogue → Engines, which is where the
accept button actually lives. The assertion follows the move; what it
pins is unchanged — the hint must name a place the user can reach it.
- Carries the CJK allowlist entries for the rendered dub bundle (#1517)
and the regenerated route snapshot for /workers/agent (#1516), both of
which this branch inherits from the workspace sync.
- docs/install/linux.md: the dictation capsule is bottom-anchored
everywhere except Wayland, where the protocol gives applications no
say in their placement. Documented rather than left as a surprise
(CodeRabbit).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: stop a flaky dependency fetch from failing green runs
en-core-web-sm resolves to a direct GitHub release URL, and github.com
intermittently answers `http2 error: refused stream before processing
any application logic`. uv's own three retries all land within the same
few seconds and fail together, so the whole job dies on a dependency
that has nothing to do with the change under test — it cost #1518 and
#1517 an otherwise-green run tonight.
Two changes: back off between whole `uv sync` attempts, which is what
actually clears it, and pass --no-sync to the pytest steps. `uv run`
re-resolves the environment before running, so every test step was a
fresh chance to hit the same fetch even though the install step had
already synced — that is exactly how #1518 failed, in the isolated
backend/tests step, with all 5467 tests already passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: one retry seam for every uv sync, not just the job that failed last
en-core-web-sm resolves to a direct GitHub *release* URL rather than a
package index, and github.com intermittently answers `http2 error:
refused stream before processing any application logic`. uv's own
retries all land inside the same ~10 seconds and fail together, so a job
dies on a dependency unrelated to the change under test. Tonight that
cost four otherwise-green runs across #1515, #1517 and #1518 — and the
first fix only covered the Tests job, so the next failure simply moved
to Smoke (Linux), which syncs separately.
The fetch is per-job, so the fix has to be per-job: scripts/uv-sync-retry.sh
backs off between whole attempts (15s, 45s, 90s) and every workflow that
syncs now goes through it — ci.yml (tests + the platform matrix),
release.yml, security.yml, evals.yml. It still fails loudly after four
attempts, so a genuinely broken lockfile is not disguised as a flake.
The Tests job also lacked the UV_HTTP_TIMEOUT / UV_HTTP_RETRIES the smoke
matrix has always set, which is part of why it was the one that kept
dying; it has them now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(ci): pin the Intel-Mac contract by intent, not by command spelling
test_ci_verifies_intel_mac_as_the_documented_remote_only_host asserted
the literal line `run: uv sync --extra pockettts`, so routing every sync
through scripts/uv-sync-retry.sh read as a broken Intel-Mac contract. The
contract it exists to protect is that the pockettts extra installs ONLY
on backend_supported legs — which the regex now pins, while leaving how
the sync is invoked free to change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* ci: keep every uv run out of the resolver, and bound the retry budget
CodeRabbit, #1517:
- `uv run` re-resolves before running, so the smoke suite, the
worker-artifact tests, the release test run and the eval run were each
a fresh chance to hit the flaky direct-URL fetch outside the retry
loop. All of them pass --no-sync now; the environment is already
synced by the step that owns the retries. security.yml's
`uv run --with pip-audit` is deliberately left alone — it layers an
ephemeral package rather than running the project's own tests.
- The retry count multiplied uv's own budget (UV_HTTP_RETRIES=5 with a
120 s timeout on the smoke matrix). Three attempts and 60 s of total
backoff outlast the refusals actually observed while staying well
inside the jobs' timeout-minutes.
- The Intel-Mac contract test pinned the smoke command literally too, so
--no-sync tripped it exactly like the sync line did. Same fix: assert
the contract (smoke runs only on backend_supported legs), not its
spelling.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
After the ordering fix, unloading the model on a 4090 still left the GPU at
1238 MiB with torch reporting 8.5 MB allocated and 803 MB reserved -- and no
number of Flush Memory presses moved it. A segment dump said why: ONE 803 MB
segment, 794.7 MB of it inactive-but-split, pinned by a single live block of
8,519,680 bytes.
That is cuBLAS's default workspace. It is taken from the caching allocator on
first use, so it lands inside whatever segment the model load had just grown,
and it is held for the life of the cuBLAS handle. empty_cache() can only
return segments that are entirely free, so one 8.5 MB block kept three
quarters of a gigabyte from ever reaching the driver again. On a machine
lending its GPU that is the difference between an idle node costing 470 MiB
and costing 1.2 GB.
free_vram() now clears the workspaces before emptying the cache, on the
unload paths only -- the next cuBLAS call re-takes one, which is cheap but
not something to pay per generate. The binding is private
(torch._C._cuda_clearCublasWorkspaces), so it is optional by construction: a
build without it keeps today's behaviour rather than failing an unload.
Found by adding reserved-vs-allocated to /system/flush-memory in 642513d2.
Allocated alone reads near zero after an unload, which is exactly why this
hid for so long -- every diagnostic we had agreed the memory was free.
The shared voice model's unload emptied the allocator caches and *then*
dropped the reference. That frees nothing: the weights are still reachable
when gc.collect() runs, empty_cache() only returns blocks the allocator
already considered free, and the reference drops a moment later into a cache
nothing will flush again. The unload logs success, the engine leaves the
registry, and nvidia-smi does not move.
Six modules open-coded the same two lines. Exactly one had them inverted --
OmniVoiceBackend.unload, which is the path the engine-registry idle sweep
reaches, which is the sweep a headless worker node runs. So every unload a
user could trigger from the UI worked, and the one that runs unattended on a
machine lending its GPU held 3.6 GB indefinitely. Found on hardware: the
sweep fired on schedule, logged "Released 1 idle engine(s)", and VRAM stayed
flat at 3656 MiB for the next two minutes.
Replace all six with model_manager.unload_shared_model(), which clears the
reference, drops the clone-prompt side cache, then frees -- in that order,
in one place. Two callers gain the side-cache drop they were missing
(/system/flush-memory and the shutdown path), which is the same defect one
step down: an unload that kept the encoded reference tensors belonging to the
model it had just released.
A source guard asserts nothing outside model_manager assigns the shared
reference, so the next caller cannot reintroduce the ordering. It caught the
sixth site while being written.
Also give the AudioSeal watermark models the bargain every other model in the
app already makes: they loaded on the first embed and stayed resident for the
life of the process. CPU-resident, so this is system RAM rather than VRAM,
and the machines that notice are the ones running batches.
The error text on a failing unload changes with the ordering. "Could not be
unloaded, retry after the current generation finishes" was accurate when the
cache flush ran first and aborted before the release; now the release has
already happened and only the flush can fail, so it says that instead of
sending the user to repeat work that is done.
The startup preload exists so the first generate feels instant for the person
sitting in front of the app. A machine lending its GPU has nobody sitting
there, so it was several GB of VRAM held from boot against a request that may
never arrive — and the idle sweep could not reclaim it, because the sweep owns
the worker executor's engines while this is the default local model.
Measured on gpu2: a node that had run nothing still sat at 2.4 GB, and an idle
unload after a real job returned it to exactly that floor rather than below it.
Worker-mode processes now load on first request and release when idle, which is
what a node should do. A machine that is both a desktop app and a worker keeps
the warm-up — there is a real user there and the point stands.
A generation abandoned at its execution budget was reported as the machine being too slow, whatever had actually happened. A job wedged on a lock it could never acquire got the same message as one genuinely grinding through a long synthesis, so the advice — use a smaller model, close other apps — was wrong exactly when the cause was a bug rather than the hardware.
The wedge decision now reads the deepest frame of the stalled thread in a stdlib module rather than pattern-matching the tail of the stack, so a short active stack cannot be mistaken for a blocked one. A thread parked in a lock, an event wait or a future is a wedge and says so; one executing engine code is slow hardware and keeps the old guidance.
A missing or ABI-mismatched torch/transformers surfaced as 'omnivoice not importable', which reads as a damaged VoiceStudio install and sent people reinstalling the app — the one thing that could not help, because the broken package is in the Python environment underneath it. The failure is now attributed to the dependency that actually failed, with the remedy that repairs it.
The second half: a model that failed to load during startup preload left the app looking healthy. Nothing surfaced the failure, so the first generation produced nothing and the cause was already gone from view. The failure and its remedy now reach the model status, where the UI can show them.
A classifier that fails while classifying no longer publishes its own raw exception text through model status.
CHANGELOG only; model_manager.py auto-merged. Also normalised #1414's
entry to the credit-then-ref order the other 42 community entries use —
CodeRabbit flagged the same inconsistency on this PR.
Two ways the repair could do the wrong gigabytes:
A shard that stays unparseable after a full re-fetch would be re-fetched
again on every generate request. The re-download now runs at most once per
repo per process, the same contract as the snapshot-link repair, and later
attempts go straight to the manual delete-and-reinstall message.
With OMNIVOICE_PRELOAD_TTS_ASR on, the load also pulls the Whisper checkpoint
— a different repo. A damaged shard there arrived looking identical, and
re-downloading the TTS checkpoint would have fixed nothing while reporting
the wrong model as broken. One local load without ASR settles which it is.
`_model_lock` is a module-level asyncio.Lock, so it binds to whichever loop first contends for it — in practice the server's. But OmniVoiceBackend._ensure_loaded() runs on a GPU-pool worker thread with no running loop and bootstraps a fresh one via asyncio.run(get_model()). Awaiting a lock owned by another loop does not block, it raises 'is bound to a different event loop', which reached users as a 500 — or deadlocks, depending on which loop touched it first.
_heal_tts_placement already carried a running_on_gpu_pool() guard for exactly this; the cold-load path never got one. It now loads inline on the calling thread.
The load must run inline rather than through _load_model_with_timeout(), which would hand _load_model_sync back to _get_gpu_pool() — the pool the caller already occupies. MPS pins that pool to a single worker, so it would wait on itself. Exclusion therefore comes from a loop-agnostic threading.Lock, not from holding a GPU slot: a slot is not exclusion when the pool has more than one worker, which CUDA hosts do.
Regression tests drive the real failure shape — a live foreign loop genuinely holding the lock, since an uncontended acquire() never binds — and install a pool that refuses submit, so a re-submission regression fails in under a second instead of hanging the suite.
An interrupted or mangled model download leaves one of two states, and only
one had any handling. A MISSING shard raises transformers' "does not appear
to have a file named …" and gets a whole recovery ladder. A shard that is
PRESENT with wrong bytes — a download stopped mid-file, a shard truncated by
antivirus, an HTML error page saved under its name — opens fine and then
fails inside safetensors:
Error while deserializing header: header too large
That reached the user as a raw 500 on every generation, from voice design and
gallery previews alike, and could not enter the ladder for two independent
reasons: the wording is not the missing-shard wording, and SafetensorError is
a Rust-extension exception rather than an OSError.
It also needs the opposite repair. The ladder RESUMES a download, and a resume
trusts a blob that is already the expected size — so it would never re-fetch
the one file that is actually wrong. The new path forces a full re-download,
then retries the load once.
Both halves now classify as MODEL_CACHE_CORRUPT: one class to the user, one
remedy, two repairs underneath. The cause is matched through the whole
exception chain, since transformers wraps the tensor library's error in its
own before it reaches us.
Two bugs from the v0.4.2 rename sweep:
- The lazy model import was rewritten to `from omnivoice.models.omnivoice import VoiceStudio`, a class the library does not export. ImportError is not ModuleNotFoundError, so the #564 source fallback never caught it and /generate 500'd on every default-engine request. The class keeps its library name — it is a checkpoint-referenced identifier, not branding.
- alembic resolved a bare relative script_location against the process cwd, and the desktop shell launches the backend from frontend/src-tauri, so a pending migration killed startup. script_location and prepend_sys_path are now anchored with %(here)s, with path_separator = os so Windows drive letters and paths with spaces survive. alembic floor raised to >=1.16.
Regression tests verified fail-before/pass-after for both.
Renames what users see. The app, the installers, the window title, the
docs and all 21 locales now say VoiceStudio, with "(previously
OmniVoice-Studio)" noted near the title of each doc surface so people
recognise it.
Deliberately NOT renamed, because renaming any of them silently breaks
an existing install — there is no legacy-path fallback anywhere in this
codebase:
- bundle identifier com.debpalash.omnivoice-studio (MSI UpgradeCode,
macOS TCC grants, managed venv, WebView localStorage, the
single-instance lock)
- data directories OmniVoice / .omnivoice and omnivoice.db
- the ~150 OMNIVOICE_* environment variables
- the X-OmniVoice-* HTTP headers (a wire protocol)
- the published Docker image paths
- the OmniVoice ENGINE, which is a model name and not this product
tests/test_identity_paths_survive_the_rename.py pins every one of those
so a future well-meaning sweep cannot orphan a user's library.
Linux .deb users install a new package name and should apt remove
omnivoice-studio; that note is in the changelog.
Two open issues arrived as a plain exit code 1 and were both handed the VRAM default: #1282 died 4s after launch on import torchaudio, and #1376 died 28s in on transformers' lazy loader raising ModuleNotFoundError. Neither user had a memory problem — their venv was half-installed, and they were sent to flush a model that had nothing to do with it.
Nothing in the exit code separates these cases; the traceback was already in the crash marker, unread because the hint only looked at exit_code/signal. It now reads the tail: an import failure naming a package the app cannot run without means the environment is incomplete, and the fix is Clean & Retry, which rebuilds it without touching voices or projects.
Only the most recent lines of the tail count, since backend_err.log is appended across runs and a process that died early carries the previous run's output. The branch sits after signal 9 (an OOM kill is an unambiguous fact about this process) and before the native-fault branch (a missing dependent DLL really does present as an access violation).
Every subprocess engine cold-loads its model inside the synthesize handler, so a first-use generate on a slow connection spent its whole 300s execution budget downloading — then failed blaming the hardware, while the sidecar's watchdog was being fed progress frames the entire time.
The outer clock now listens to that evidence: SubprocessBackend forwards each sidecar progress frame to a per-worker-thread heartbeat (pool jobs only, cleared when the job ends), and the guarded waiter extends the deadline past the soft budget only while heartbeats stay fresher than MODEL_LOAD_HEARTBEAT_GRACE_S, bounded by MODEL_LOAD_EXTRA_TIMEOUT_S. The extension is logged once. A silent job still dies at the original deadline; stopped heartbeats kill within the grace; another thread's heartbeat is no alibi; caller cancellation cancels and consumes the abandoned future. 11 regression tests, headline case verified failing before.
* fix(gpu): record WHERE a wedged job was stuck when its budget expires
Three reporters on v0.4.2 hit "TTS generate ran for more than 300s of
actual compute time and was abandoned" (#1338, #1329, #1348) — two of them
on an RTX 3050 and an RTX 3060, rendering a single sentence. That is not a
machine too slow for the job. The message says "too heavy for the
available compute" because it is the only story the timeout path can tell.
And nothing in the log could contradict it. The timeout branch logged THAT
the budget was exceeded, reset the pool, and returned. The worker cannot
be cancelled, so it was still running on a real stack — and we threw that
away, which is why every report of this class arrives undiagnosable and
the only advice available is "reproduce it under a debugger".
sys._current_frames() reads the frame of every live thread including one
wedged inside a C call, which is exactly this case. Filtered to gpu-pool
workers so the log names the stuck job rather than the web server, capped
at 25 frames, and it can never raise — a diagnostic that throws would
replace a real GpuJobTimeoutError with an unrelated crash.
Ordering is load-bearing and asserted: the capture runs BEFORE reset(),
because reset() swaps in a fresh executor and the wedged thread then stops
being identifiable as a pool worker — the diagnostic would still run, still
log, and be empty, which looks like it worked.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(gpu): redact home paths from captured stacks; label stale workers
Both review findings on this PR, both valid.
CodeRabbit (CWE-532): traceback frames carry absolute source paths, which
on a user machine start with their home directory — their account name.
This log lands in backend.log, which goes into diagnostic bundles and
prefilled bug reports, so it has to be sanitized like every other surfaced
text. Reuses core.failure.sanitize rather than inventing a second answer;
if sanitizing itself fails the stacks are dropped, not logged raw.
greptile P1: a wedged worker survives reset() — it cannot be cancelled and
keeps running under the same gpu-pool name the replacement pool uses. The
second timeout in a session would log both with nothing to tell them
apart, and the stale one is the more misleading, since it names an
operation that is not the job that just failed. The live pool is now
identified through its own thread set and the others are marked STALE.
That set comes from ThreadPoolExecutor._threads, which is private, so
unknown internals degrade to labelling nothing rather than to failing —
a diagnostic that vanishes because an attribute moved is worse than an
unlabelled one, and that degrade path has its own test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The report arrived as a raw traceback pasted out of a log file — because that is the only place the reason existed. A failed chapter was a red row and the word "failed"; the SSE event carried the literal string "chapter failed to render", the symptom the user could already see.
- Both the per-chapter and the terminal all-failed events are now built with the shared `core.failure` builder: sanitized text, a guaranteed non-empty reason, error class, docs deeplink and hint. `error` keeps mirroring `reason`, so older frontends and the Stories exporter are unaffected. The chapter list shows the reason inline, full text in the row tooltip.
- **A silent infinite hang on the same path.** asyncio refuses to put `StopIteration` into a Future — `_copy_future_state` raises TypeError inside the event loop's own callback, so the `run_in_executor` future is never completed and the caller waits forever, with no error, event or timeout. Reachable from ordinary input: VoxCPM's `next_and_close` is a bare `next(gen)`, so a generator ending without yielding raises it straight into a GPU-pool worker. Fixed at the pool boundary (`_ResilientGpuPool.submit` plus a matching `_cpu_pool` subclass) as `WorkerStopIteration(RuntimeError)`, keeping the original as `__cause__`.
- An all-failed render now marks the job failed. It previously returned without touching job history, so the row stayed `running` and the next startup's orphan sweep read a hopeless render as interrupted and offered it as resumable (Greptile P1).
Tests: 6 pool-guard tests (two hang without the fix, bounded so a hang fails rather than stalls CI), longform e2e extended for both event shapes, the empty-`str(exc)` floor, and both sides of the job-status branch, plus the chapter-list component.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(engines): make SubprocessBackend.generate() path-aware on the GPU pool
generate()'s slot handling had two bugs:
1. On-pool self-deadlock: /v1/audio/speech and /generate (and audiobook,
dub, batch) dispatch generate() via run_on_gpu_pool_guarded, already on
a pool worker, so the inner slot submit queued behind the very job
running it on a 1-worker (MPS) pool and timed out before the sidecar
spawned. Every subprocess engine surfaced the in-process 300s-abandon
instead of synthesizing.
2. Off-pool no hold: the off-pool slot was a bare no-op that released the
worker before _spawn(), so off-pool callers (engine self-test,
diagnostics) could synthesize concurrently with a pool job and
over-subscribe the GPU.
Make the slot block path-aware: on-pool callers skip (the outer
run_on_gpu_pool_guarded already holds _running for the whole sidecar
exchange); off-pool callers hold a real slot for the whole synthesis via
an _occupy task that blocks the worker until _held is set in the finally.
Single release point in the finally.
Regression tests: generate dispatched on a pool worker (on-pool skip) and
a concurrent pool job blocked during an off-pool generate (off-pool hold).
Both verified fail-before / pass-after.
Supersedes #1296 (on-pool-skip-only). Closes#1295, #1297.
* Address review: couple on-pool skip to the pool prefix; fix comment
/simplify + /code-review flagged that the on-pool skip keyed on the literal
"gpu-pool" string, decoupled from _build_gpu_pool's thread_name_prefix. A
rename would silently re-introduce the exact self-deadlock this PR fixes (and
the tests can't catch it, since they hardcode the prefix). Centralise the
prefix in _GPU_POOL_THREAD_PREFIX + a running_on_gpu_pool() helper, used by
_build_gpu_pool, the skip in generate(), and _heal_tts_placement.
Also fix the comment: the Settings engine self-test rejects subprocess-isolated
engines with a 400, so the only real off-pool caller is the diagnose.py
deep-synth probe.
* fix(engines): bind slot_future before the off-pool branch
CodeQL py/uninitialized-local-variable (error, blocking CI). `_held is not
None` does imply slot_future was assigned, so the current code is correct —
but the two are only coupled by convention, which the analyser cannot see and
a third exit path would quietly break. Binds it to None up front and guards
the cancel.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(engines): make the slot-hold regression deterministic and leak-free
CodeRabbit, valid on both counts. The test used sleep(0.8)/sleep(0.5) as
synchronization — the tests/** contract forbids it, and on a slow runner the
marker could be enqueued before the generator had reserved anything, so the
assertion passed for the wrong reason. It now waits on an event signalled when
the slot task actually starts, and asserts "did not run" via a result()
timeout rather than a bare sleep.
Cleanup moved into finally: an assertion failure used to leak the sidecar
process and the pool thread into the rest of the session.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(engines): add omnivoice-subprocess, a crash-isolated TTS engine
The default in-process OmniVoice engine runs on the GPU ThreadPoolExecutor.
When a generate or load exceeds its execution budget the pool is "reset", but
the abandoned worker thread cannot be killed (Python cannot interrupt a native
torch/MPS call), so it keeps holding the device until it finishes on its own
and later synths queue behind it and hang. The reset restores pool capacity
but not the device. This is the residual root cause behind the closed#730
and #1190: the messaging/reset mitigations address the symptom, not the
device-holding zombie.
Add an opt-in `omnivoice-subprocess` engine that runs the same model in a
child process via SubprocessBackend. A child process can be hard-killed: on a
recv-timeout the watchdog calls proc.kill(), reclaiming VRAM/device, and the
next request transparently respawns a fresh sidecar. The in-process engine
remains the default, so existing users see no change; this is an opt-in for
unattended / scheduled / reaction-triggered synthesis where a stuck job must
self-recover instead of hanging until a manual restart.
Base-class and mitigation changes that ship with it:
- SubprocessBackend.generate() now consumes non-terminal {"op":"progress"}
frames a sidecar emits during a cold load (previously the first cold
generate after spawn failed, then worked on retry). Additive: engines that
reply with audio directly are unaffected.
- recv_timeout_s is overridable per engine (default 60s unchanged); the new
engine sets it to the generate budget so a long-but-valid synth is not
falsely killed while a wedged one still is.
- make_room_before_generate(): free idle GPU memory before a warm, heavy
generate. The cold-load path already evicted; the warm path skipped it, so a
long synth on a VRAM-tight MPS box could contend its way into the budget.
Verified end-to-end against the live model (cold / warm / recovery-after-kill)
and under a sustained + concurrent-pressure soak: killed-worker recovery 5/5,
chunked long text 9/9, no memory leak.
* Address review: install_hint + move make_room into get_model
- Add `omnivoice-subprocess` to `_INSTALL_HINTS`; the
test_install_hints_cover_all_registered_backends gate requires every
registered backend to carry one (this was the CI failure).
- Move the warm-generate VRAM eviction out of the /generate and
/v1/audio/speech routes and into get_model()'s warm-return path, so EVERY
native TTS generate is covered (REST, WS TTS, dub, batch, audiobook), not
just the two REST routes. Drops the now-redundant per-route wiring.
(Greptile P1: the per-route placement missed the other generation surfaces.)
* Address review: drop dead long-text eviction path; log probe failure
- _should_make_room_for_generate: the long-text headroom boost became dead
code once the eviction moved into get_model() (which has no text), so the
long-text branch never fired. Removed the text param, the long-text
threshold/multiplier branch, and the now-unused _env_float helper. The core
RAM-tight gate (the part that matters on a starved box) is unchanged.
- Log the available_memory probe failure at debug instead of silently
swallowing it (CodeRabbit: silent swallow breaks the debug trail).
- Tests updated for the text-agnostic policy.
* fix(engines): stop subprocess generate() self-deadlock on 1-worker pools
SubprocessBackend.generate() acquires a GPU-pool slot for accounting, but
/v1/audio/speech and /generate dispatch backend.generate() via
run_on_gpu_pool_guarded, i.e. already ON a pool worker. On a 1-worker pool
(MPS) the inner pool.submit queued behind the very job running it and
slot_future.result(timeout=10) raised before the sidecar ever spawned, so
omnivoice-subprocess (and every other subprocess engine on MPS) surfaced the
in-process 300s-abandon instead of synthesizing.
Skip the slot acquisition when current_thread() is already a gpu-pool worker;
the outer guard already accounts for the slot. Direct callers (off the pool)
still acquire one. Regression test added (generate on a pool worker).
* Address review: reword slot-skip comment (fixes watermark-coverage CI) + simplify
- The slot-skip comment said "dispatch backend.generate() via", and
test_watermark_route_coverage's _SYNTH_CALL regex matches the literal
backend.generate( anywhere in a module, so it counted subprocess_backend.py
as a synthesis producer that must reference mark_synthetic (it doesn't — the
routes apply mark_synthetic; the engine sits below the chokepoint, like
tts_backend.py). Reworded to "dispatch generate() via".
- Fold in the simplify refinement: single negated predicate, import+pool
moved into the acquire branch.
* fix(engines): warn about under-provisioned hardware before the synth, not after
Six reports are the same story: #1240, #1246, #1248, #1277, #1283, #1284 —
4 GB and 6 GB cards running an engine that wants 6 GB, each one waiting out
the full 300s compute budget to be told the job "was too heavy". The routing
layer knew the whole time. The error text even names the card and the figure.
The caveat only ever surfaced on the engine-PICK toast, so it reached people
who changed engines and nobody whose engine was already selected — the
default, or one persisted from a previous session. That is most users.
/generate does return X-OmniVoice-Routing, but a response header arrives when
the job ends, five minutes too late to be a warning.
So the check moves to the chokepoint every synth path shares (api/generate.ts,
same argument as the in-flight count). Fire-and-forget: never awaited, so it
cannot add latency to the request it warns about; never throws, so an
unreachable backend costs a warning rather than a generate; once per
engine+reason per session, so it informs instead of nagging. Advisory, not
blocking — the driver can page to system RAM and short inputs fit where long
ones don't.
Extracts routingNotice() as the single frontend mirror of the backend's
routing_notice(). Two callers now need "is this verdict worth interrupting
for", and two inline copies would drift — invisibly, until someone on DirectML
or an unavailable engine gets a hardware warning for a normal pick.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(cuda): stop sending every RTX 40-series card to the CPU
The SM-arch gate required the device's exact tag in get_arch_list(). NVIDIA's
rules are not exact, and PyTorch depends on that: SASS is binary-compatible
UPWARD within a major version, so the official wheels ship sm_80/sm_86 and
deliberately no sm_89 — the 8.6 kernels already cover Ada. Exact matching
therefore declared sm_89 unsupported, check_device_compatibility() returned
False, and get_best_device() silently returned "cpu".
That is every RTX 4060/4070/4080/4090, not just the reporter's card (#1285) —
each one running TTS on the CPU on hardware that works fine, with a message
telling them their GPU was unsupported.
cuda_build_covers() now applies the real rules: sm_XY covers same-major
devices with minor >= Y; compute_XY PTX JITs forward to anything newer; an
a/f suffix (sm_90a) is architecture-specific and matches exactly. Unparseable
entries are skipped, and an empty arch list still degrades to "compatible" —
the pre-existing fail-open contract.
The remediation text also pointed at a NIGHTLY index for what is a stable
supported card; it now names the stable cu128 index.
12 tests: the Ada regression, Jetson Orin (8.7), downward-within-major and
cross-major rejection, PTX forward-JIT, arch-specific suffixes, and a genuine
sm_120-on-old-wheel mismatch so the gate is proven to still work.
Closes#1285
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(cuda): resolve app modules at call time, not import time
The tests/** review contract forbids module-level imports of app modules —
they go stale under sys.modules pollution from other suites, which is the live
cause of #1269's cross-suite failures. Binds core.device_caps per call.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: restore generate.ts and generatePreflight.test.js from main
Conflict markers were committed in the previous merge — `git add` on the
directory staged both files as resolved while the markers were still in them.
Both belong to #1288 and are unchanged by this PR, so they take main's version
verbatim.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(dub): purge order was hash-dependent, and the cap could evict a live marker
main went red on my own test. Two distinct defects, both mine.
1. `targets = set(job_ids)` made iteration order depend on PYTHONHASHSEED, so
which markers a cap-forced trim discarded was luck. That is why the test
passed locally and failed in CI — verified: the old code passes at seeds
0/7/42 and fails at 12345. Now a de-duplicated list in caller order.
2. The size cap could evict markers the CURRENT purge had just recorded. Those
are the newest and the likeliest to still be held by a running job, so
dropping one is precisely the resurrection this mechanism exists to prevent.
A 'clear history' larger than the cap forced exactly that. The cap now never
touches the current purge, making the real bound cap + one purge — stated
plainly rather than implied.
Tests now pin both: identical survivors across runs, and an oversized purge
keeping all of its own markers. Verified across six hash seeds; full suite green
under 12345, the seed that reddened main.
* feat(update): announce updates as a toast with actions, not a wall of text
A version's release notes ARE the whole changelog section — v0.4.1's was 42
bullets. Any surface that renders them inline becomes unusable: older builds put
them in a blocking OS dialog that filled the screen and had to be dismissed
before the app could be touched.
Removing that dialog left the opposite failure. The only remaining signal was a
6-pixel dot beside the version number in the footer, which is easy to never
notice — so users either got shouted at or told nothing.
A toast is the middle: it names the version, offers Install and restart /
What's new / Later, and leaves. The notes stay one click away in Settings →
Updates, where there is room. Keyed by version so the 6-hourly re-check
replaces rather than stacks, and it never auto-dismisses — an update the user
hasn't answered is still true. Install declines while a generation is running,
since the relaunch would lose it.
Strings added to all 21 locales, translated rather than English-filled.
Tests pin the shape that matters: the toast takes no notes prop at all, renders
under 200 characters, and cannot stack duplicates.
* fix(update): don't relaunch while work is in flight; share the busy check
Review of #1272 found the restart guard was `dubStep === 'generating'` and
nothing else. Installing an update relaunches the process, so that permitted
throwing away a dub upload, a transcription, a translation, an export or a
standalone TTS synth. The same narrow check was written twice — in the new
toast and in UpdatesPanel — so the two could also drift apart.
Replaced with a single `isAppBusy(state)` in utils/appBusy, unioning the
signals the store actually has: the dub state machine, the floating status
pill (which every long background operation already pushes to), and
`ttsGenerating` — a new transient store field mirroring useTTS's local
`isGenerating`, which lived in a hook where no global check could see it.
`dubStep === 'editing'` is deliberately not busy: it waits on the user, and
counting it would block updates for as long as a transcript stays open.
Also from review:
- A failed lazy import of the toast fell into the outer catch, whose
setUpdateIdle() erased the update that had just been found. The
announcement is optional; the available state is not.
- "Dismiss" was machine-translated into the employment sense — terminate an
employee — in de/ja/ru/zh-CN/zh-TW, on close buttons and one aria-label.
Swept every locale rather than the four lines that were flagged.
- Hoisted the test's mock state with vi.hoisted.
CodeRabbit's changelog finding is declined: Highlights bullets carry no issue
refs by design, which tests/test_changelog_style.py enforces.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(shutdown): a quit mid-generate is not a 500, and not a bug report (#1276)
#1174 made a model load interrupted by shutdown benign for the background
preload, but a *request* that triggered a load took the generic unhandled-
exception path: crash log, ERROR traceback, and an error-journal entry that
feeds the bug-report pipeline. Quitting the app with a generate queued
surfaced "500 Internal Server Error: model load skipped: backend shutting
down" and offered to file a GitHub issue for a normal teardown.
Nothing failed — the process is exiting. The handler now answers 503 with
Retry-After and an actionable detail, ahead of the crash-log/journal writes.
Frontend half of the same bug: toastErrorWithReport offered "Report" for any
error. A 503 means "not now, try again" by definition, so it now shows the
backend's message without the report action — covering a still-warming
backend too, not just this shutdown path.
Fail-before/pass-after tests on both sides, including that the shutdown case
leaves no crash-log entry and no journal record.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(models): repair a half-downloaded model whichever way it reports (#1273)
transformers has two unrelated wordings for "this snapshot has no weight
shard", sharing no words:
hub load "<repo> does not appear to have a file named …"
local dir "Error no file named model.safetensors, … found in directory …"
The self-heal and the failure classifier both matched only the first. The
second is what a load of a *subfolder* inside a cached snapshot raises —
exactly where an interrupted download leaves a half-written repo — so the
reporter got neither the automatic repair (delete broken entries →
re-download → retry) nor an actionable hint, just a raw 500. Their disk had
10.6 GB free, i.e. a download that ran out of room.
The phrase list now lives once in core.failure, so the healer and the error
text cannot disagree about what an interrupted download looks like. Both
fragments of the second wording must match — "no file named" alone is
ordinary English and must not claim the class.
Also hardens the #1276 handler found by running both suites in one session:
services.model_manager can be imported under two module names, making two
distinct ModelLoadInterruptedByShutdown classes and breaking a bare
isinstance — which silently restored the 500 this fix exists to remove. Now
matched by isinstance OR class name, with a test that raises a same-named
class from a different module.
Verified against the full tests/ + backend/tests/ session: the only remaining
failures are the four pre-existing #1269 isolation leaks, unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(update): count synths in flight at the chokepoint, not in one caller
Review round 2. Both Greptile P1s were the same weakness in my first pass:
tracking the synth in useTTS meant only the Generate tab counted (voice
previews, the compare modal, the stories editor and profile previews call
generateSpeech directly and were invisible), and a boolean meant two
overlapping syntheses cleared each other — whichever settled first reported
"idle" while the other was still running, so Install and restart discarded it.
Moved to `api/generate.ts`, around the one `/generate` call all seven paths
share, as a count with a `finally` release (so an abort or a network error
frees it too). A future synth caller is covered without opting in.
Fail-before verified: the overlap test fails against the boolean version.
Also from review:
- UpdatesPanel re-reads the busy state at click time; the render-time snapshot
only exists to disable the button, and work can start after the last render.
- The 503 now carries the allowed-origin CORS headers. Without them the
browser reports a bare CORS failure and the actionable detail — the whole
point of the fix — never reaches the user. Both error responses build them
through one helper now.
- Log `request.url.path`, not the full URL: a query string can carry tokens
and newlines.
- `update.busy` said "finish your dub first" in all 21 locales, but the guard
now covers uploads, transcription, translation, export and synth. Rewritten
as work-in-progress wording, translated per locale.
- ru `common.dismissStatus` was "reject status"; missed in the earlier sweep.
Declined: CodeRabbit's request for a `(#NNN)` ref on the Highlights bullet —
Highlights carry no refs by design (CLAUDE.md), and tests/test_changelog_style.py
enforces it. Also declined gating the cache re-download behind a confirmation:
the self-heal already existed and already ran for the sibling wording, this
only stops it missing half the class, and it stays behind the existing
_hf_offline() check.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(update): grey out Install while work is running
CI lint caught `busy` as unused after the click-time check moved into the
handler — and that exposed a real gap: `busy` had never been wired to
anything. The comment claimed it disabled the button; it didn't. Clicking
Install during a synth just bounced a toast back.
Now it does what it said: the Install and Restart buttons are disabled while
work is in flight, with `update.busy` as the tooltip. The click-time read
stays the authority, since work can start between the last render and the
click — the disabled state is the explanation, not the safety.
Adds the first UpdatesPanel test. The case it pins hardest is the inverse of
the bug: a busy predicate stuck at true would make the app permanently
un-updatable, which is worse than what this fixes. So it asserts enabled when
idle and during dub 'editing' (which waits on the user), disabled during an
upload, a translation, and one or more synths.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
GpuJobTimeoutError was raised inside `except asyncio.TimeoutError` without
`from`, so the original timeout was dropped from the traceback.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two P1s, both correct:
- Being IN the override map was treated as proof of compatibility. If the
wheel ships neither the native arch nor the remap target, setting
HSA_OVERRIDE_GFX_VERSION only changes WHICH kernel is missing — gfx1151 with
a gfx1030-only build was routed to the GPU and would fail at launch. Both
arch_unsupported() and _configure_rocm_if_needed() now require the target to
be present, and fall back to CPU otherwise.
- An EMPTY arch list means the build's metadata is unavailable, not that the
GPU is unsupported. The remap branch read that unknown state as a confirmed
mismatch and would push a natively-supported gfx1151 onto foreign gfx1100
kernels. Now fails open and changes nothing, matching the fail-open contract
the rest of the probe follows.
ROCM_GFX_OVERRIDES values are now the target gfx NAME rather than the HSA
version string, so the "is the target present?" check is a direct membership
test; hsa_override_for() derives the env-var form, covered by a test that
every entry in the map converts cleanly.
Also: CHANGELOG entries shortened with refs last, and the MD028 blank line
between the two docker.md blockquotes (CodeRabbit).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two P1s, both correct:
- MPS mislabel. HostCaps.vram_gb on MPS is a heuristic (system RAM / 2) for a
UNIFIED memory pool, so an 8 GB Mac reports 4.0 "VRAM" — comparing that to a
floor measured on discrete CUDA hardware would warn every small Mac about an
engine that runs fine there. The caveat is now dedicated-VRAM families only
(cuda/rocm); MPS has a different memory model and no measured floor.
- Engine-agnostic timeout. _timeout_guidance serves EVERY job on the GPU pool
(reference transcribe, stream assemble, watermarking, dub steps, CPU-only
engines on a GPU host), and a hardcoded 6 GB threshold applied without
knowing whose job it is would confidently misdiagnose most of them. The
floor is now passed in via run_on_gpu_pool_guarded, defaulting to 0 — so the
under-provisioned wording is opt-in and only the TTS generate dispatches opt
in. A test asserts every "TTS generate" dispatch passes it, so the branch
can't become unreachable in production.
CHANGELOG entries reworded to end with their refs (CodeRabbit).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>