d91beef0fd314250d8d9b94de86dfea019a8bd96
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
58c6f37252 |
perf(dub): single-use per-segment refs no longer evict the prompts a dub reuses; add docs/performance.md (#1132)
* perf(dub): single-use per-segment refs no longer evict the prompts a dub reuses; add docs/performance.md The scan-resistance fix: A dub cuts a distinct reference clip per segment (Wave 3.2 / #486 — each line clones its own source delivery) and falls back to the per-speaker clone for segments under 3 s. Both paths flow through the voice-clone prompt cache — an LRU of 8. Streaming hundreds of one-shot per-segment clips through that LRU evicts the per-speaker and locked-profile prompts that every fallback segment reuses, so the speaker ref was re-encoded (~0.4 s each, measured with scripts/bench_pipeline.py) again and again across the render. Note what this deliberately does NOT do: the bench's "166 misses vs 2 speakers" framing suggested keying refs per speaker — but per-segment refs are the intentional prosody-matching feature, and the re-transcription behind them is the #1004 correctness fix. Their encode cost is the price of the feature, not waste. The waste was only the eviction side-effect, and that's what this removes: _get_clone_prompt(store=False) still reads the cache (a hit is free) but never inserts, and the dub loop marks exactly the segment-scoped refs (auto-seg: bindings and auto: bindings resolved to a segment clip) as single-use. Per-speaker, locked-profile, and preview refs cache as before. cache_ref is popped in generate_with_cached_ref before the model call — the model's generate() has an explicit signature and would TypeError — and unknown engines ignore it (**kw adapters). The doc: docs/performance.md is the first performance documentation in the repo — none of the ~15 perf env vars appeared anywhere in docs/, the Performance panel's only control is Windows-only, and slowness reports (#1032) arrived as mysteries instead of settings checks. Covers the three classic causes of "it got slow", where generation/dub time goes, every knob with defaults and warnings (raising OMNIVOICE_GPU_WORKERS on a small GPU is the #567 crash, not a speedup), platform notes, and how to run the bench so reports carry numbers. Linked from README's install section. Tests: store=False semantics (encodes, never inserts, still reads), the flood scenario end to end (a speaker prompt stays warm through 3x the cache cap of one-shots), and the pop contract (cache_ref never reaches the model). Full suite: 2974 passed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs,dub: review round — qualify the per-file cache claim; note the OOM-retry tradeoff - CodeRabbit: docs/performance.md's "the reference encode is cached per file" now carves out the dub's per-line clips (single-use by design — nothing for a cache to save). - Greptile P2 (OOM retry re-encodes a single-use ref): acknowledged in a code comment as deliberate — caching the retry's ref would reintroduce the eviction this flag prevents, to optimize a path that only runs after an OOM already cost seconds. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(performance): probe-based torch.compile wording; honest accelerator + cache claims (review) Greptile's repeated OOM-retry finding is deliberately skipped: retaining the prompt across the retry would require passing prompt objects through the adapter protocol (backend.generate takes paths), to save 0.4s on a path that only runs after an OOM already cost seconds — the tradeoff is documented at the call site. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c613a65435 |
perf(tts,dub): the reference clip was re-encoded on every chunk; the dub loaded a 3 GB model to throw it away (#1130)
* perf(tts,dub): the reference clip was re-encoded on every chunk; the dub loaded a 3 GB model to throw it away Two independent pieces of pure waste on the generate path, both measured with scripts/bench_pipeline.py on a 16 GB M2 (a reference encode costs 0.40s). 1. The voice-clone prompt cache was built, then orphaned. #427/#473 added a bounded LRU that encodes a reference clip once and reuses it, because "every cloned generation re-encodes the reference audio from scratch". It was wired into OmniVoiceBackend — the *adapter* path. But /generate for the default engine forks to the *native* model path (that fork predates the cache, #324) and passed ref_audio=<path> straight through, so the codec encoder re-ran the reference on every model.generate() call: once per text chunk, once per pause-span, once per audiobook segment, and once per request. That perf PR has therefore only ever sped up /v1/audio/speech. The Generate button never touched it. Every native call site now goes through one helper (generate_with_cached_ref) so the rule lives in a single place: chunked /generate, its streaming twin (#1088), the [pause] stitcher (#276), and the audiobook renderer. Saving is 0.40s x (calls - 1): ~3.6s on a 10-chunk text, ~66s on a 166-segment audiobook. Same-class bug found in the same cache: /v1/audio/speech accepts preprocess_prompt, but the adapter dropped it before it reached the model AND the cache key omitted it — so the flag was silently ignored, and honoring it without keying on it would have served (and poisoned) the wrong prompt. Both fixed together. 2. A dub loaded the TTS core just to free it again. The transcribe preflight called get_model() — pulling in the ~3 GB TTS model — for one reason: to read a preloaded `_asr_pipe` off it. That attribute only exists under OMNIVOICE_PRELOAD_TTS_ASR, which is off by default. So every dub loaded the model, harvested None, had offload_tts_for_asr() free it 60 lines later (on unified memory that is a full UNLOAD, #1119), and then cold-reloaded the same model in dub_generate (~8s). Load -> unload -> reload, for an attribute that was always None. It now loads only when there is something to harvest. Also fixes a latent NameError: asr_on_vocals was assigned only inside the model-loaded branch but read from _gen_body, so an early preflight bail raised NameError instead of the real error. Tests: the existing cache tests passed the whole time the cache was dead, because they test the cache in isolation with a stub model. The new tests assert the wiring instead — that a real render encodes the reference ONCE regardless of how many generate calls it takes. All four encode-count tests fail before this change and pass after; the dub tests likewise. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(changelog): the reference re-encode and the dub's throwaway model load (#1130) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(memory): unloading the TTS model must drop its cached reference prompts too Follow-on to the cache wiring in this PR, and a real gap it opened. clear_clone_prompt_cache() was called from exactly one place: OmniVoiceBackend.unload(). That was sufficient while the prompt cache was adapter-only — but the native /generate path now populates it, and the native path unloads through model_manager (idle_worker, _offload_unified_memory), not through the adapter. So cached prompt tensors would have survived an unload. That directly undercuts #1119: on unified memory offload_tts_for_asr() sets model = None precisely to hand the RAM to the ASR model. Prompts left behind sit in the memory the unload was trying to reclaim. The tensors are small (integer codes, not waveforms), so this is hygiene rather than a leak — but "unload means unload" is the whole point of that change, and the next thing cached here might not be small. model_manager.release_tts_side_caches() is now called wherever the global model is dropped. Best-effort by construction: cache hygiene must never be able to break an unload, because a failed unload is how the backend gets OOM-killed. The test binds services.tts_backend at CALL time, not import time: several suites purge sys.modules["services.*"] for DB isolation (test_model_load_timeout, test_model_manager_preload), so a module-level alias goes stale mid-run and the assertion would inspect a different module's cache than the code under test just filled. Production already imports it at call time. Full suite: 2968 passed, in both deterministic and random order. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): keep the prompt cache best-effort, and stop the unload hook closing an import cycle Three review findings, all real. 1. Greptile P1 — the shared helper dropped the inline fallback. OmniVoiceBackend.generate() has always caught a failure from generate(voice_clone_prompt=...) and retried with the inline ref, so the cache stays a pure latency optimization. generate_with_cached_ref did not: a model that rejected a precomputed prompt would have turned a working /generate, streaming render, or audiobook job into a hard error. Moving the native path onto the cache would then have made it LESS robust than before it was cached at all. The helper now carries that fallback, and OmniVoiceBackend delegates to it instead of keeping a second copy. Two subtly-diverging copies of this logic is precisely how the cache ended up wired into the adapter and nowhere else; there is now exactly one. 2. CodeQL — cyclic import. release_tts_side_caches() imported services.tts_backend, which already imports model_manager: a real cycle, not a false positive. A registration hook fixed the cycle but replaced it with a worse problem — the hook runs at import time and pulls model_manager (and core.config) in earlier than before, which perturbs DATA_DIR binding and broke test_longform_jobs::test_route_handler_returns_jobs_envelope in the full suite (passed in isolation, failed in order — caught locally, not in CI). It now reaches the module through sys.modules instead: no import, no cycle, no import-time side effect. And it is the more correct expression of the invariant anyway — a module that was never imported has no cache to clear. 3. CodeRabbit — the audiobook and streaming call sites had no encode-count test. Added one for the audiobook synth path (the worst case: hundreds of segments on one voice). New tests fail before their respective fixes: stripping the try/except from the helper fails the prompt-rejection test. Full suite: 2970 passed, deterministic and random order. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: mergetest <nizam4103@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4a0d18f510 |
perf(omnivoice): cache voice-clone prompt embeddings (#427) (#473)
Every cloned generation re-encoded the reference audio from scratch — a fixed per-request latency that compounds on batch / long-form / dataset workloads that reuse one saved voice across many calls. The OmniVoice model already exposes the fast path (create_voice_clone_prompt → VoiceClonePrompt, generate(voice_clone_prompt=)); the Studio backend just wasn't using it. OmniVoiceBackend.generate now: - builds a VoiceClonePrompt once per reference and caches it (bounded LRU, max 8, keyed by ref path + mtime + ref_text; thread-safe — generation runs in a GPU thread pool), then passes voice_clone_prompt= to skip the re-encode; - falls back to the inline ref_audio/ref_text path on ANY cache miss or error, so output is identical either way (the model documents the two as equivalent) — this is purely a latency optimization, never a behaviour change; - the design/instruct path (no ref_audio) is untouched. - unload() clears the cache so a flush / engine-switch frees the prompt tensors. tests/test_clone_prompt_cache.py: 6 cases (encode-once-then-hit, ref_text + mtime invalidation, LRU eviction at the cap, encode-failure → None fallback, clear). 6 passed. Closes #427. |