Commit Graph
1 Commits
Author SHA1 Message Date
cbcea41fb6 feat(memory): honest /model/loaded accounting + a free-memory budget probe (#1111)
Two gaps the model-management investigation surfaced, now closed.

1. /model/loaded reported only the OmniVoice core, so a resident second engine
   (mlx-audio, cosyvoice, …) and the warm dictation ASR were INVISIBLE — the
   memory picture looked ~2 GB lighter than reality on exactly the boxes that
   OOM. list_loaded() now enumerates the in-process engine instances (from the
   generate path's cache) and the capture ASR singleton too, and adds a
   `system` block: free/total RAM (and free VRAM on a dedicated GPU) plus a
   low-memory advisory. Verified live: after an mlx-audio generate the panel
   shows `engine:mlx-audio` and `system: {ram_available_gb, ram_total_gb}`,
   where before it showed nothing.

2. services/memory_budget.py: available_memory() reads FREE memory now (device
   caps only reports total, once per process) — free system RAM via psutil,
   free VRAM via torch.cuda.mem_get_info on a dedicated GPU; on MPS the RAM
   figure is what matters (unified memory). low_memory_warning() returns an
   advisory below a headroom threshold (OMNIVOICE_LOW_MEMORY_HEADROOM_GB,
   default 2). The generate path calls log_if_low() before a load, so a later
   OOM kill leaves a breadcrumb pointing at the load that tipped it instead of
   a silent death.

Advisory only — nothing is blocked: the OS reclaims cache, and refusing a load
on an estimate would brick machines that would cope. The single-active-engine
eviction (#1105) is what actually reclaims room; this makes the picture honest
and leaves forensics.

6 new unit tests (threshold logic / VRAM-precedence / never-raises); frontend
LoadedModelsResponse typed for the new `system` field + id shapes. Backend
suite 2918 passed; typecheck clean.

Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 16:11:42 +05:30