Commit Graph
1 Commits
Author SHA1 Message Date
Palash Debnathandmergetest e9477870c8 fix(memory): on unified memory, "offload" must mean UNLOAD — the 16 GB dub OOM (#1119) (#1122)
offload_tts_for_asr() exists to make room before WhisperX large-v3 (~3 GB) loads
for a dub. On CUDA it moves the TTS model to CPU. On Apple Silicon it did
NOTHING — an early return with the comment "MPS / CPU / DirectML don't benefit
from manual offloading".

That reasoning is right about the STRATEGY and wrong about the CONCLUSION. On
unified memory, moving a model "to CPU" frees nothing, because it is the same
physical RAM. But that means the fix is to RELEASE the model — not to skip
making room altogether.

Measured on a 16 GB M2, at the moment a dub begins:
    TTS model resident      3,107 MB
    backend footprint       4,170 MB
    free RAM                 4.17 GB
large-v3 then wants ~3 GB of that, alongside the app and macOS. The OS kills the
backend mid-transcription, and the stream "drops before emitting any segments".

On a unified-memory host the TTS model is now actually released when free RAM is
below a headroom threshold (default 6 GB, OMNIVOICE_UNIFIED_OFFLOAD_HEADROOM_GB),
and left warm when there's room — so a roomy machine pays no reload. get_model()
lazily reloads it on the next generation, so restore is correctly a no-op. The
CUDA path is untouched.

Verified end to end in a real process: model loaded → offload_tts_for_asr() →
`mm.model is None` and free RAM recovered. Previously it returned immediately and
freed nothing.

6 tests (releases when tight / stays warm when roomy / restore is a no-op /
no model is a no-op / a failing probe never aborts the dub / CUDA path unchanged).

This is a CAUSE, not another error-message fix.

Refs #1119 #1113

Co-authored-by: mergetest <nizam4103@gmail.com>
2026-07-12 18:33:36 +05:30