d91beef0fd314250d8d9b94de86dfea019a8bd96
1
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bb813ff676 |
feat(startup): bind the socket in ~1s and narrate startup step by step (#1550)
* feat(startup): bind the socket in ~1s and narrate startup step by step The structural fix for the "can't reach the local backend" class (~1 in 5 of every issue ever filed): uvicorn served nothing until torch import (10-20s cold), the 30-router fan-out, an import-time DB migration, the cuDNN preload, and alembic all finished — every slow or fragile step rendered as an unexplained dead backend. main.py now keeps module scope fast and defers the heavy work: - _phase_a_build (executor thread): prefs/env restore + #963 migration, yt-dlp overlay, cuDNN preload, torchaudio, model_manager, router imports — order preserved, literal imports so PyInstaller still traces. - _phase_a_finalize (event loop, no awaits → atomic wrt requests): include_router, mounts, MCP, SPA, openapi bust. - _phase_b: the old lifespan startup body; handles on app.state so shutdown survives a startup that never finished. - Eager mode (pytest / OMNIVOICE_EAGER_INIT=1) runs everything at import — byte-equivalent behavior for the ~100 lifespan-less TestClient sites and for embedders (dump_api_routes, probe boot runner opt in). While starting: /health answers 503 with the current step, new /startup/progress serves the full ledger (always 200), and StartupGateMiddleware 503s everything else with the [starting] marker (same skip-the-Report-button convention as [shutting_down]). A deferred failure keeps import-crash semantics: traceback to stderr → shell crash forensics, run sentinel stays uncleared, exit 1 names the failed step. Shell: startup_progress() probe (marker-header-gated so a foreign responder can't narrate the splash) feeds per-step log lines into the launch poll and the supervisor's reconnect wait. --health-check absorbs the deferred init (60→180s); --diagnose runs Phase A up front so it still sees restored prefs. Docker HEALTHCHECK semantics unchanged (curl -f fails on 503 exactly as it did on connection-refused). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(startup): join the Phase A thread on shutdown; async fail-path sleep Bot-review harvest on #1550: cancelling the deferred-startup task cannot stop the executor thread inside Phase A's blocking imports — shutdown now waits (bounded, only when a build started and hasn't finished) on a thread-completion event so interpreter teardown can't race a mid-import (#1000 class). The failure path's last-poll beat is now awaited, not time.sleep — a blocking sleep froze the very loop that beat exists to let serve. Also: dump_api_routes forces eager (assignment, not setdefault), and the integration test's child gets DEVNULL instead of an undrained pipe that could wedge a cold boot. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(startup): close the Phase A submission race; CodeQL nits Review finds on #1550: shutdown could sample _phase_a_started unset while the executor callable was queued-but-not-running, skipping the thread join. started is now set BEFORE submission, the submission is shielded so a cancel can't strand a queued callable that would never set _phase_a_finished, and the wrapper sets finished on every exit including the already-built early return. Contract pinned by test_phase_a_thread_join_contract. Plus explanatory comments on the new bare excepts and a consistent return in the gate's websocket branch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |