Surfaces a routing verdict for the CURRENTLY-SELECTED TTS engine in the two
system-health surfaces, so a CPU fallback / unavailable-GPU is heard about
before a slow or failed synth — the no-silent-fallback contract, read-only.
- `tts_backend.active_routing()` + `gpu_routing_verdict()`: the active engine's
routing derived from list_backends() (byte-identical to the matrix) plus the
host compute summary (family + VRAM from the canonical probe). Never raise.
- `/system/diagnose` gains a `gpu_routing` check: accelerated→ok,
accelerated-with-caveat / cpu_fallback→warn (+ actionable hint), cpu_only→ok
(no-GPU host is the expected normal state — never noise-warns), unavailable→
fail, no-engine→warn. ASCII-safe detail strings (the text dump enforces ASCII).
- `/setup/preflight` gains an "Active engine routing" check + an explicit
`gpu_routing` object on PreflightResponse (a real field — the response has no
extra="allow", so it would otherwise be dropped). `device` gains `gpu_family`
(ROCm-vs-CUDA aware) + `vram_gb`. New `GpuRouting` schema.
Tests: gpu_routing_verdict (host + active-engine + degraded), diagnose status
mapping across all 6 states + never-raises, preflight gpu_routing object +
check + device.gpu_family. Existing diagnose/preflight tests stay green (checks
are additive; the report's top-level key set is unchanged).
Deferred (documented): synth-time routing headers/WS-frames at the 3 synth
entry points. Selection is already hard-gated (PR 3 select_engine), and the
matrix (PR 5) + this preflight/diagnose verdict surface the situation — the
synth-time signal is incremental belt-and-suspenders for the env-var-pinned
edge and is best validated interactively. Tracked as a #21 follow-up.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>