Files
VoiceStudio/backend/services/engine_memory.py
T
8750b91522 feat(memory): one TTS engine resident at a time — stop stacking models on 16 GB (#1105)
Measured on a 16 GB M2: a generate on omnivoice (~2.8 GB core) followed by a
generate on mlx-audio left BOTH resident (footprint 3.9 → 4.3 GB) because the
OmniVoice core lives in model_manager.model while every other engine lives in
engines._ENGINE_INSTANCES — two caches that never coordinated, and the latter
was never unloaded. That accumulation is a direct contributor to the memory
pressure behind the "Can't reach the local backend" OOM deaths.

- services/engine_memory.py: evict_other_tts_engines(keep_id) unloads every
  OTHER resident TTS engine before the incoming one loads — spans both stores
  (the OmniVoice core under its async lock, and the instance cache). No-op when
  nothing else is resident, so steady-state single-engine use pays nothing; only
  a real switch evicts. Default on; OMNIVOICE_SINGLE_ENGINE_RESIDENT=0 to keep
  several warm. Wired into the /generate path right after the engine resolves.
- TTSBackend.unload() (the ABC default) now actually frees the held model: it
  clears _MODEL_ATTRS (_model/_tts) and empties the device cache. Every
  in-process engine but OmniVoice previously inherited a NO-OP unload(), so an
  engine switch dropped the instance ref but left its model for GC with the GPU
  cache un-emptied. One change fixes all of them and is future-proof.
- FasterWhisperBackend.unload() cleared self._asr — an attribute it never
  assigns — so its model in self._model was never freed. Fixed.

Live-verified: omnivoice → mlx-audio now DROPS footprint 2300 → 1541 MB (core
evicted) instead of climbing to 4305 with both resident. 7 new unit tests
(eviction spans both stores / keeps the active engine / no-op when disabled /
a failing unload doesn't abort the sweep / ABC unload frees + is idempotent),
order-independent. Backend suite green.

Refs the 16 GB OOM class (#1076 #1092 #1093 #1101)

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 15:05:11 +05:30

100 lines
4.1 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Single-active-TTS-engine memory discipline.
Only one TTS engine's model stays resident at a time. When the generate path
resolves an engine, every *other* resident engine is unloaded first — so the
previous engine's model is handed back instead of stacking in memory until GC.
Why this matters (measured on a 16 GB M2): a generate on ``omnivoice`` leaves
its ~2.8 GB core model resident; a subsequent generate on ``mlx-audio`` loaded
that engine's model **on top** (footprint 3.9 GB → 4.3 GB, both resident),
because the two live in different caches with no coordination — the core in
``model_manager.model``, the rest in ``engines._ENGINE_INSTANCES`` (which was
never unloaded). That accumulation is the baseline that pushes a 16 GB machine
into the memory pressure behind the "Can't reach the local backend" OOM deaths.
Default on. Opt out with ``OMNIVOICE_SINGLE_ENGINE_RESIDENT=0`` on machines with
RAM to spare (keeping several engines warm avoids the reload latency on an A/B
switch — ~8 s for the OmniVoice core, ~1–2 s for the lighter engines).
"""
from __future__ import annotations
import logging
import os
logger = logging.getLogger("omnivoice.engine_memory")
_OFF = {"0", "false", "no", "off"}
def single_engine_resident() -> bool:
"""Whether the one-engine-at-a-time policy is active (default True)."""
return (os.environ.get("OMNIVOICE_SINGLE_ENGINE_RESIDENT", "1").strip().lower()
not in _OFF)
def _evict_instance_cache(keep_cls) -> list[str]:
"""Unload + drop every cached engine instance except ``keep_cls``.
Operates on the per-request instance cache the generate path shares with the
engine health route (``engines._ENGINE_INSTANCES``). Each engine's
``unload()`` frees its heavy model (the ABC default clears ``_MODEL_ATTRS``
and empties the device cache; subprocess engines reap their sidecar). Never
raises — a stuck unload must not block the generation that triggered it."""
evicted: list[str] = []
try:
from api.routers.engines import _ENGINE_INSTANCES
except Exception: # pragma: no cover — router import should always succeed
return evicted
for cls, inst in list(_ENGINE_INSTANCES.items()):
if cls is keep_cls:
continue
try:
inst.unload()
except Exception: # noqa: BLE001
logger.warning("evict: %s.unload() failed", getattr(cls, "id", cls.__name__),
exc_info=True)
_ENGINE_INSTANCES.pop(cls, None)
evicted.append(getattr(cls, "id", cls.__name__))
return evicted
async def evict_other_tts_engines(keep_id: str) -> list[str]:
"""Unload every resident TTS engine except ``keep_id`` and return their ids.
Spans both stores a TTS model can live in: the OmniVoice core singleton
(``model_manager.model``, freed under its async lock when we're switching
*away* from it) and the generic engine instance cache. A no-op when the
policy is off or nothing else is resident, so steady-state single-engine use
pays nothing — only an actual switch evicts. Never raises."""
if not single_engine_resident():
return []
evicted: list[str] = []
# The OmniVoice core singleton — only when the incoming engine isn't it.
if keep_id != "omnivoice":
try:
import services.model_manager as mm
async with mm._model_lock:
if mm.model is not None:
mm.model = None
mm.free_vram()
evicted.append("omnivoice")
except Exception: # noqa: BLE001
logger.warning("evict: OmniVoice core unload failed", exc_info=True)
# Every other in-process / sidecar engine instance.
keep_cls = None
try:
from services.tts_backend import get_backend_class
keep_cls = get_backend_class(keep_id)
except Exception: # noqa: BLE001 — unknown id → evict all cached instances
keep_cls = None
evicted.extend(_evict_instance_cache(keep_cls))
if evicted:
logger.info("single-engine eviction: freed %s (keeping %s)", evicted, keep_id)
return evicted