* fix(engines): warn before a long CPU synth burns the whole budget
#1288 closed the under-provisioned-GPU gap but left the CPU one open, and I
missed it: a CPU-only host is a BENIGN routing verdict, so routingNotice()
correctly stays silent — yet #1299 and #1260 are exactly that shape, CPU hosts
that hit the 300s budget on long text with no warning at all. "Nothing is
misconfigured" and "this will finish in time" are different claims.
Threshold is the backend's own definition of past-short: generate_timeout_for()
gives the first 1200 characters the flat budget before extending it, so
ordinary sentences on a CPU laptop stay quiet and only the shape that actually
times out is flagged. Hardware caveats still take precedence — one toast, and
it names the real reason rather than generic advice.
5 tests; engines.cpuLongText translated in all 21 locales.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(engines): don't tell CPU-tuned engines to switch to themselves
Greptile P1. The advice names OmniVoice GGUF and Supertonic-3 as the CPU-tuned
alternatives — shown to someone already running one of them, it is advice to
switch to what they are using. Those two now get the same warning without the
self-referential clause; the engine set matches the backend's own timeout
message so the two can't disagree about who is CPU-tuned.
Also documents the preflight in docs/performance.md (docs-sync rule): both
warning shapes, why the threshold is 1200 characters (it is the figure the
budget itself uses), that they are advisory and once-per-engine-per-session,
and the CPU-tuned exception.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
A job queued behind a busy 1-worker pool burned its whole 300s budget
without executing an instruction, then reported "too heavy for the
available compute". The clock now starts when a worker picks the job up;
queue wait has its own generous bound and surfaces as a retryable
saturation error.
Also: reset() no longer cancels innocent queued peers; the timeout
message stops claiming capacity was restored (the abandoned job keeps the
device until it drains); every GPU dispatch uses the shared length-scaled
budget; watermark embeds move off the GPU pool; /v1/audio/speech gets
429/503 + Retry-After; a timed-out batch segment fails the job instead of
shipping a silent gap.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`offload_tts_for_asr()` moves the TTS model to CPU to make VRAM room for
WhisperX, but its partner `restore_tts_after_asr()` was only reachable on
the dub-transcribe success path. Any abort, terminal error, or client
disconnect skipped it, and `get_model()` never re-checked placement — so
EVERY subsequent /generate ran on CPU (10-50x slower, CPU pegged) until
the ~15-minute idle unload happened to fire. Reported as "speed varies by
time of day"; it is fully deterministic.
Two independent guarantees:
- Balance the pair at the call site: gen()'s `finally` now pays the
restore debt on every exit path, chained off the ASR unload so the two
never contend for VRAM (and fire-and-forget, since the finally also
runs under GeneratorExit where awaiting is illegal).
- Self-heal placement (the class fix): `get_model()` verifies the model
is on the resolved target device and moves it back if not, so a future
unbalanced offload path cannot strand it either. Cheapest-first probe —
one parameter check on the hot path; unified memory is exempt (its
offload releases the model rather than moving it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Only the TTS model (~2.4 GB) is required on first run; ASR models are
per-platform curated picks (curated_on in models.yaml) installed on demand.
Every transcription surface returns a typed asr_model_missing error with a
one-click download CTA instead of silently pulling multi-GB Whisper weights.
Settings -> Models is a grouped, platform-aware catalog. New guided
permissions UX (wizard System Check + Settings -> Permissions + mic
pre-flight) with native mic-state checks and OS settings deep-links. New
parakeet-mlx engine brings Parakeet TDT v3 to Apple Silicon (language-gated
capture preference so multilingual dictation never regresses). Docs:
expressive-speech page, Flush/Unload + CPU-fallback triage, clone-length FAQ.
Hardening: preflight fails open for custom model pins, ROCm curation no
longer inherits NVIDIA picks, Windows mic probe reads the NonPackaged
consent key, CaptureWidget setup race fixed, offline-cache CI simulation
fixes so empty-cache runners stay green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* perf(dub): single-use per-segment refs no longer evict the prompts a dub reuses; add docs/performance.md
The scan-resistance fix:
A dub cuts a distinct reference clip per segment (Wave 3.2 / #486 — each line
clones its own source delivery) and falls back to the per-speaker clone for
segments under 3 s. Both paths flow through the voice-clone prompt cache — an
LRU of 8. Streaming hundreds of one-shot per-segment clips through that LRU
evicts the per-speaker and locked-profile prompts that every fallback segment
reuses, so the speaker ref was re-encoded (~0.4 s each, measured with
scripts/bench_pipeline.py) again and again across the render.
Note what this deliberately does NOT do: the bench's "166 misses vs 2 speakers"
framing suggested keying refs per speaker — but per-segment refs are the
intentional prosody-matching feature, and the re-transcription behind them is
the #1004 correctness fix. Their encode cost is the price of the feature, not
waste. The waste was only the eviction side-effect, and that's what this
removes: _get_clone_prompt(store=False) still reads the cache (a hit is free)
but never inserts, and the dub loop marks exactly the segment-scoped refs
(auto-seg: bindings and auto: bindings resolved to a segment clip) as
single-use. Per-speaker, locked-profile, and preview refs cache as before.
cache_ref is popped in generate_with_cached_ref before the model call — the
model's generate() has an explicit signature and would TypeError — and unknown
engines ignore it (**kw adapters).
The doc:
docs/performance.md is the first performance documentation in the repo — none
of the ~15 perf env vars appeared anywhere in docs/, the Performance panel's
only control is Windows-only, and slowness reports (#1032) arrived as mysteries
instead of settings checks. Covers the three classic causes of "it got slow",
where generation/dub time goes, every knob with defaults and warnings (raising
OMNIVOICE_GPU_WORKERS on a small GPU is the #567 crash, not a speedup), platform
notes, and how to run the bench so reports carry numbers. Linked from README's
install section.
Tests: store=False semantics (encodes, never inserts, still reads), the flood
scenario end to end (a speaker prompt stays warm through 3x the cache cap of
one-shots), and the pop contract (cache_ref never reaches the model). Full
suite: 2974 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs,dub: review round — qualify the per-file cache claim; note the OOM-retry tradeoff
- CodeRabbit: docs/performance.md's "the reference encode is cached per file"
now carves out the dub's per-line clips (single-use by design — nothing for
a cache to save).
- Greptile P2 (OOM retry re-encodes a single-use ref): acknowledged in a code
comment as deliberate — caching the retry's ref would reintroduce the
eviction this flag prevents, to optimize a path that only runs after an OOM
already cost seconds.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(performance): probe-based torch.compile wording; honest accelerator + cache claims (review)
Greptile's repeated OOM-retry finding is deliberately skipped: retaining the
prompt across the retry would require passing prompt objects through the
adapter protocol (backend.generate takes paths), to save 0.4s on a path that
only runs after an OOM already cost seconds — the tradeoff is documented at
the call site.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>