offload_tts_for_asr() exists to make room before WhisperX large-v3 (~3 GB) loads
for a dub. On CUDA it moves the TTS model to CPU. On Apple Silicon it did
NOTHING — an early return with the comment "MPS / CPU / DirectML don't benefit
from manual offloading".
That reasoning is right about the STRATEGY and wrong about the CONCLUSION. On
unified memory, moving a model "to CPU" frees nothing, because it is the same
physical RAM. But that means the fix is to RELEASE the model — not to skip
making room altogether.
Measured on a 16 GB M2, at the moment a dub begins:
TTS model resident 3,107 MB
backend footprint 4,170 MB
free RAM 4.17 GB
large-v3 then wants ~3 GB of that, alongside the app and macOS. The OS kills the
backend mid-transcription, and the stream "drops before emitting any segments".
On a unified-memory host the TTS model is now actually released when free RAM is
below a headroom threshold (default 6 GB, OMNIVOICE_UNIFIED_OFFLOAD_HEADROOM_GB),
and left warm when there's room — so a roomy machine pays no reload. get_model()
lazily reloads it on the next generation, so restore is correctly a no-op. The
CUDA path is untouched.
Verified end to end in a real process: model loaded → offload_tts_for_asr() →
`mm.model is None` and free RAM recovered. Previously it returned immediately and
freed nothing.
6 tests (releases when tight / stays warm when roomy / restore is a no-op /
no model is a no-op / a failing probe never aborts the dub / CUDA path unchanged).
This is a CAUSE, not another error-message fix.
Refs #1119#1113
Co-authored-by: mergetest <nizam4103@gmail.com>