`offload_tts_for_asr()` moves the TTS model to CPU to make VRAM room for
WhisperX, but its partner `restore_tts_after_asr()` was only reachable on
the dub-transcribe success path. Any abort, terminal error, or client
disconnect skipped it, and `get_model()` never re-checked placement — so
EVERY subsequent /generate ran on CPU (10-50x slower, CPU pegged) until
the ~15-minute idle unload happened to fire. Reported as "speed varies by
time of day"; it is fully deterministic.
Two independent guarantees:
- Balance the pair at the call site: gen()'s `finally` now pays the
restore debt on every exit path, chained off the ASR unload so the two
never contend for VRAM (and fire-and-forget, since the finally also
runs under GeneratorExit where awaiting is illegal).
- Self-heal placement (the class fix): `get_model()` verifies the model
is on the resolved target device and moves it back if not, so a future
unbalanced offload path cannot strand it either. Cheapest-first probe —
one parameter check on the hot path; unified memory is exempt (its
offload releases the model rather than moving it).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>