Commit Graph
3 Commits
Author SHA1 Message Date
58c6f37252 perf(dub): single-use per-segment refs no longer evict the prompts a dub reuses; add docs/performance.md (#1132)
* perf(dub): single-use per-segment refs no longer evict the prompts a dub reuses; add docs/performance.md

The scan-resistance fix:

A dub cuts a distinct reference clip per segment (Wave 3.2 / #486 — each line
clones its own source delivery) and falls back to the per-speaker clone for
segments under 3 s. Both paths flow through the voice-clone prompt cache — an
LRU of 8. Streaming hundreds of one-shot per-segment clips through that LRU
evicts the per-speaker and locked-profile prompts that every fallback segment
reuses, so the speaker ref was re-encoded (~0.4 s each, measured with
scripts/bench_pipeline.py) again and again across the render.

Note what this deliberately does NOT do: the bench's "166 misses vs 2 speakers"
framing suggested keying refs per speaker — but per-segment refs are the
intentional prosody-matching feature, and the re-transcription behind them is
the #1004 correctness fix. Their encode cost is the price of the feature, not
waste. The waste was only the eviction side-effect, and that's what this
removes: _get_clone_prompt(store=False) still reads the cache (a hit is free)
but never inserts, and the dub loop marks exactly the segment-scoped refs
(auto-seg: bindings and auto: bindings resolved to a segment clip) as
single-use. Per-speaker, locked-profile, and preview refs cache as before.

cache_ref is popped in generate_with_cached_ref before the model call — the
model's generate() has an explicit signature and would TypeError — and unknown
engines ignore it (**kw adapters).

The doc:

docs/performance.md is the first performance documentation in the repo — none
of the ~15 perf env vars appeared anywhere in docs/, the Performance panel's
only control is Windows-only, and slowness reports (#1032) arrived as mysteries
instead of settings checks. Covers the three classic causes of "it got slow",
where generation/dub time goes, every knob with defaults and warnings (raising
OMNIVOICE_GPU_WORKERS on a small GPU is the #567 crash, not a speedup), platform
notes, and how to run the bench so reports carry numbers. Linked from README's
install section.

Tests: store=False semantics (encodes, never inserts, still reads), the flood
scenario end to end (a speaker prompt stays warm through 3x the cache cap of
one-shots), and the pop contract (cache_ref never reaches the model). Full
suite: 2974 passed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs,dub: review round — qualify the per-file cache claim; note the OOM-retry tradeoff

- CodeRabbit: docs/performance.md's "the reference encode is cached per file"
  now carves out the dub's per-line clips (single-use by design — nothing for
  a cache to save).
- Greptile P2 (OOM retry re-encodes a single-use ref): acknowledged in a code
  comment as deliberate — caching the retry's ref would reintroduce the
  eviction this flag prevents, to optimize a path that only runs after an OOM
  already cost seconds.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(performance): probe-based torch.compile wording; honest accelerator + cache claims (review)

Greptile's repeated OOM-retry finding is deliberately skipped: retaining the
prompt across the retry would require passing prompt objects through the
adapter protocol (backend.generate takes paths), to save 0.4s on a path that
only runs after an OOM already cost seconds — the tradeoff is documented at
the call site.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 14:36:06 +05:30
c613a65435 perf(tts,dub): the reference clip was re-encoded on every chunk; the dub loaded a 3 GB model to throw it away (#1130)
* perf(tts,dub): the reference clip was re-encoded on every chunk; the dub loaded a 3 GB model to throw it away

Two independent pieces of pure waste on the generate path, both measured with
scripts/bench_pipeline.py on a 16 GB M2 (a reference encode costs 0.40s).

1. The voice-clone prompt cache was built, then orphaned.

#427/#473 added a bounded LRU that encodes a reference clip once and reuses it,
because "every cloned generation re-encodes the reference audio from scratch".
It was wired into OmniVoiceBackend — the *adapter* path. But /generate for the
default engine forks to the *native* model path (that fork predates the cache,
#324) and passed ref_audio=<path> straight through, so the codec encoder re-ran
the reference on every model.generate() call: once per text chunk, once per
pause-span, once per audiobook segment, and once per request.

That perf PR has therefore only ever sped up /v1/audio/speech. The Generate
button never touched it.

Every native call site now goes through one helper (generate_with_cached_ref) so
the rule lives in a single place: chunked /generate, its streaming twin (#1088),
the [pause] stitcher (#276), and the audiobook renderer. Saving is
0.40s x (calls - 1): ~3.6s on a 10-chunk text, ~66s on a 166-segment audiobook.

Same-class bug found in the same cache: /v1/audio/speech accepts
preprocess_prompt, but the adapter dropped it before it reached the model AND
the cache key omitted it — so the flag was silently ignored, and honoring it
without keying on it would have served (and poisoned) the wrong prompt. Both
fixed together.

2. A dub loaded the TTS core just to free it again.

The transcribe preflight called get_model() — pulling in the ~3 GB TTS model —
for one reason: to read a preloaded `_asr_pipe` off it. That attribute only
exists under OMNIVOICE_PRELOAD_TTS_ASR, which is off by default. So every dub
loaded the model, harvested None, had offload_tts_for_asr() free it 60 lines
later (on unified memory that is a full UNLOAD, #1119), and then cold-reloaded
the same model in dub_generate (~8s). Load -> unload -> reload, for an attribute
that was always None. It now loads only when there is something to harvest.

Also fixes a latent NameError: asr_on_vocals was assigned only inside the
model-loaded branch but read from _gen_body, so an early preflight bail raised
NameError instead of the real error.

Tests: the existing cache tests passed the whole time the cache was dead, because
they test the cache in isolation with a stub model. The new tests assert the
wiring instead — that a real render encodes the reference ONCE regardless of how
many generate calls it takes. All four encode-count tests fail before this change
and pass after; the dub tests likewise.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(changelog): the reference re-encode and the dub's throwaway model load (#1130)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(memory): unloading the TTS model must drop its cached reference prompts too

Follow-on to the cache wiring in this PR, and a real gap it opened.

clear_clone_prompt_cache() was called from exactly one place:
OmniVoiceBackend.unload(). That was sufficient while the prompt cache was
adapter-only — but the native /generate path now populates it, and the native
path unloads through model_manager (idle_worker, _offload_unified_memory), not
through the adapter. So cached prompt tensors would have survived an unload.

That directly undercuts #1119: on unified memory offload_tts_for_asr() sets
model = None precisely to hand the RAM to the ASR model. Prompts left behind sit
in the memory the unload was trying to reclaim. The tensors are small (integer
codes, not waveforms), so this is hygiene rather than a leak — but "unload means
unload" is the whole point of that change, and the next thing cached here might
not be small.

model_manager.release_tts_side_caches() is now called wherever the global model
is dropped. Best-effort by construction: cache hygiene must never be able to
break an unload, because a failed unload is how the backend gets OOM-killed.

The test binds services.tts_backend at CALL time, not import time: several suites
purge sys.modules["services.*"] for DB isolation (test_model_load_timeout,
test_model_manager_preload), so a module-level alias goes stale mid-run and the
assertion would inspect a different module's cache than the code under test just
filled. Production already imports it at call time.

Full suite: 2968 passed, in both deterministic and random order.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tts): keep the prompt cache best-effort, and stop the unload hook closing an import cycle

Three review findings, all real.

1. Greptile P1 — the shared helper dropped the inline fallback.

OmniVoiceBackend.generate() has always caught a failure from
generate(voice_clone_prompt=...) and retried with the inline ref, so the cache
stays a pure latency optimization. generate_with_cached_ref did not: a model that
rejected a precomputed prompt would have turned a working /generate, streaming
render, or audiobook job into a hard error. Moving the native path onto the cache
would then have made it LESS robust than before it was cached at all.

The helper now carries that fallback, and OmniVoiceBackend delegates to it
instead of keeping a second copy. Two subtly-diverging copies of this logic is
precisely how the cache ended up wired into the adapter and nowhere else; there
is now exactly one.

2. CodeQL — cyclic import.

release_tts_side_caches() imported services.tts_backend, which already imports
model_manager: a real cycle, not a false positive. A registration hook fixed the
cycle but replaced it with a worse problem — the hook runs at import time and
pulls model_manager (and core.config) in earlier than before, which perturbs
DATA_DIR binding and broke test_longform_jobs::test_route_handler_returns_jobs_envelope
in the full suite (passed in isolation, failed in order — caught locally, not in CI).

It now reaches the module through sys.modules instead: no import, no cycle, no
import-time side effect. And it is the more correct expression of the invariant
anyway — a module that was never imported has no cache to clear.

3. CodeRabbit — the audiobook and streaming call sites had no encode-count test.
Added one for the audiobook synth path (the worst case: hundreds of segments on
one voice).

New tests fail before their respective fixes: stripping the try/except from the
helper fails the prompt-rejection test.

Full suite: 2970 passed, deterministic and random order.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 12:55:32 +05:30
Palash Debnath 4a0d18f510 perf(omnivoice): cache voice-clone prompt embeddings (#427) (#473)
Every cloned generation re-encoded the reference audio from scratch — a fixed
per-request latency that compounds on batch / long-form / dataset workloads that
reuse one saved voice across many calls.

The OmniVoice model already exposes the fast path (create_voice_clone_prompt →
VoiceClonePrompt, generate(voice_clone_prompt=)); the Studio backend just wasn't
using it. OmniVoiceBackend.generate now:
- builds a VoiceClonePrompt once per reference and caches it (bounded LRU, max 8,
  keyed by ref path + mtime + ref_text; thread-safe — generation runs in a GPU
  thread pool), then passes voice_clone_prompt= to skip the re-encode;
- falls back to the inline ref_audio/ref_text path on ANY cache miss or error,
  so output is identical either way (the model documents the two as equivalent)
  — this is purely a latency optimization, never a behaviour change;
- the design/instruct path (no ref_audio) is untouched.
- unload() clears the cache so a flush / engine-switch frees the prompt tensors.

tests/test_clone_prompt_cache.py: 6 cases (encode-once-then-hit, ref_text +
mtime invalidation, LRU eviction at the cap, encode-failure → None fallback,
clear). 6 passed.

Closes #427.
2026-06-14 22:08:35 +05:30