Files
VoiceStudio/tests/test_model_load_timeout.py
T
Palash DebnathandClaude Opus 4.8 5a723c0408 fix(model): bound first-run model load/download so it never hangs forever (#173)
Windows users reported 'create demo voice runs indefinitely, no audio, no
error'. Root cause: the first /generate triggers OmniVoice.from_pretrained()
which downloads multi-GB weights via the legacy LFS path (HF_HUB_DISABLE_XET=1),
with NO timeout anywhere. A stalled socket (proxy/firewall/AV) blocks the GPU-
pool worker forever inside get_model() -- before the try/except that would
surface an error -- and the frontend /generate fetch had no abort, so the
spinner spun forever with no toast.

- backend/main.py: set HF_HUB_ETAG_TIMEOUT=15 + HF_HUB_DOWNLOAD_TIMEOUT=30
  (per-read timeout: resets on each chunk, so slow-but-progressing downloads
  are never punished; only a dead socket trips it). Set before hf import.
- model_manager: get_model()/preload_model() now load via _load_model_with_timeout(),
  an asyncio.wait_for backstop (OMNIVOICE_MODEL_LOAD_TIMEOUT, default 1200s) that
  drops the poisoned GPU pool and raises a clear, actionable RuntimeError so a
  retry gets a fresh worker instead of queueing behind the wedged one.
- useTTS.js: AbortController backstop on /generate so the UI never spins forever
  even if the backend is unreachable; friendly timeout toast.
- tests: watchdog raises + resets pool + releases lock; env/floor parsing.

Cross-platform (no OS-specific behavior); backward-compatible with installed
models; local-first preserved.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 20:50:21 +05:30

64 lines
2.3 KiB
Python

"""Regression: model load/download must never hang forever (Windows demo-voice
"runs indefinitely, no audio, no error" report).
`get_model()` wraps the blocking load in an `asyncio.wait_for` deadline; on
timeout it drops the poisoned GPU pool and raises a clear RuntimeError instead
of leaving `/generate` (and the UI spinner) wedged forever.
"""
from __future__ import annotations
import asyncio
import sys
import threading
import pytest
@pytest.fixture
def model_manager(monkeypatch):
for mod_name in ("core.config", "services.model_manager"):
if getattr(sys.modules.get(mod_name), "__file__", None) is None:
sys.modules.pop(mod_name, None)
import services.model_manager as mm
monkeypatch.setattr(mm, "model", None)
return mm
def test_model_load_timeout_respects_env(model_manager, monkeypatch):
mm = model_manager
monkeypatch.delenv("OMNIVOICE_MODEL_LOAD_TIMEOUT", raising=False)
assert mm._model_load_timeout() == 1200.0
monkeypatch.setenv("OMNIVOICE_MODEL_LOAD_TIMEOUT", "5000")
assert mm._model_load_timeout() == 5000.0
monkeypatch.setenv("OMNIVOICE_MODEL_LOAD_TIMEOUT", "not-a-number")
assert mm._model_load_timeout() == 1200.0
monkeypatch.setenv("OMNIVOICE_MODEL_LOAD_TIMEOUT", "1") # below the safety floor
assert mm._model_load_timeout() == 30.0
def test_get_model_times_out_and_resets_pool(model_manager, monkeypatch):
mm = model_manager
# Isolated lock so we never reuse one bound to a previous test's loop.
monkeypatch.setattr(mm, "_model_lock", asyncio.Lock())
monkeypatch.setattr(mm, "_model_load_timeout", lambda: 0.3)
release = threading.Event()
def _hang(): # simulates a wedged download that never returns
release.wait(2.0)
return object()
monkeypatch.setattr(mm, "_load_model_sync", _hang)
assert mm._get_gpu_pool() is not None # a pool exists before the timeout
try:
with pytest.raises(RuntimeError, match="timed out"):
asyncio.run(mm.get_model())
assert mm.model is None # no half-loaded model
assert mm._gpu_pool_singleton is None # poisoned pool was dropped
assert not mm._model_lock.locked() # lock released for a retry
finally:
release.set() # let the orphaned worker exit immediately