Two gaps the model-management investigation surfaced, now closed.
1. /model/loaded reported only the OmniVoice core, so a resident second engine
(mlx-audio, cosyvoice, …) and the warm dictation ASR were INVISIBLE — the
memory picture looked ~2 GB lighter than reality on exactly the boxes that
OOM. list_loaded() now enumerates the in-process engine instances (from the
generate path's cache) and the capture ASR singleton too, and adds a
`system` block: free/total RAM (and free VRAM on a dedicated GPU) plus a
low-memory advisory. Verified live: after an mlx-audio generate the panel
shows `engine:mlx-audio` and `system: {ram_available_gb, ram_total_gb}`,
where before it showed nothing.
2. services/memory_budget.py: available_memory() reads FREE memory now (device
caps only reports total, once per process) — free system RAM via psutil,
free VRAM via torch.cuda.mem_get_info on a dedicated GPU; on MPS the RAM
figure is what matters (unified memory). low_memory_warning() returns an
advisory below a headroom threshold (OMNIVOICE_LOW_MEMORY_HEADROOM_GB,
default 2). The generate path calls log_if_low() before a load, so a later
OOM kill leaves a breadcrumb pointing at the load that tipped it instead of
a silent death.
Advisory only — nothing is blocked: the OS reclaims cache, and refusing a load
on an estimate would brick machines that would cope. The single-active-engine
eviction (#1105) is what actually reclaims room; this makes the picture honest
and leaves forensics.
6 new unit tests (threshold logic / VRAM-precedence / never-raises); frontend
LoadedModelsResponse typed for the new `system` field + id shapes. Backend
suite 2918 passed; typecheck clean.
Co-authored-by: mergetest <nizam4103@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>