* feat(startup): bind the socket in ~1s and narrate startup step by step The structural fix for the "can't reach the local backend" class (~1 in 5 of every issue ever filed): uvicorn served nothing until torch import (10-20s cold), the 30-router fan-out, an import-time DB migration, the cuDNN preload, and alembic all finished — every slow or fragile step rendered as an unexplained dead backend. main.py now keeps module scope fast and defers the heavy work: - _phase_a_build (executor thread): prefs/env restore + #963 migration, yt-dlp overlay, cuDNN preload, torchaudio, model_manager, router imports — order preserved, literal imports so PyInstaller still traces. - _phase_a_finalize (event loop, no awaits → atomic wrt requests): include_router, mounts, MCP, SPA, openapi bust. - _phase_b: the old lifespan startup body; handles on app.state so shutdown survives a startup that never finished. - Eager mode (pytest / OMNIVOICE_EAGER_INIT=1) runs everything at import — byte-equivalent behavior for the ~100 lifespan-less TestClient sites and for embedders (dump_api_routes, probe boot runner opt in). While starting: /health answers 503 with the current step, new /startup/progress serves the full ledger (always 200), and StartupGateMiddleware 503s everything else with the [starting] marker (same skip-the-Report-button convention as [shutting_down]). A deferred failure keeps import-crash semantics: traceback to stderr → shell crash forensics, run sentinel stays uncleared, exit 1 names the failed step. Shell: startup_progress() probe (marker-header-gated so a foreign responder can't narrate the splash) feeds per-step log lines into the launch poll and the supervisor's reconnect wait. --health-check absorbs the deferred init (60→180s); --diagnose runs Phase A up front so it still sees restored prefs. Docker HEALTHCHECK semantics unchanged (curl -f fails on 503 exactly as it did on connection-refused). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(startup): join the Phase A thread on shutdown; async fail-path sleep Bot-review harvest on #1550: cancelling the deferred-startup task cannot stop the executor thread inside Phase A's blocking imports — shutdown now waits (bounded, only when a build started and hasn't finished) on a thread-completion event so interpreter teardown can't race a mid-import (#1000 class). The failure path's last-poll beat is now awaited, not time.sleep — a blocking sleep froze the very loop that beat exists to let serve. Also: dump_api_routes forces eager (assignment, not setdefault), and the integration test's child gets DEVNULL instead of an undrained pipe that could wedge a cold boot. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(startup): close the Phase A submission race; CodeQL nits Review finds on #1550: shutdown could sample _phase_a_started unset while the executor callable was queued-but-not-running, skipping the thread join. started is now set BEFORE submission, the submission is shielded so a cancel can't strand a queued callable that would never set _phase_a_finished, and the wrapper sets finished on every exit including the already-built early return. Contract pinned by test_phase_a_thread_join_contract. Plus explanatory comments on the new bare excepts and a consistent return in the gate's websocket branch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
126 lines
3.8 KiB
Python
126 lines
3.8 KiB
Python
"""Startup progress ledger — what the backend is doing before it can serve.
|
|
|
|
Why this exists: the project's #1 lifetime failure class is "can't reach the
|
|
local backend", and a large slice of it was never a dead backend at all —
|
|
just one that couldn't say "I'm starting, currently loading PyTorch" because
|
|
nothing listened until every heavy import and migration finished. main.py now
|
|
binds the socket early and defers the heavy work; this module is the shared
|
|
state the early `/health` + `/startup/progress` endpoints report from while
|
|
that work runs.
|
|
|
|
Thread-safety: the deferred init runs Phase A in an executor thread while the
|
|
event loop serves probes, so every mutation and snapshot takes the lock.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import threading
|
|
import time
|
|
|
|
# Execution order matters only for display; the ledger records whatever order
|
|
# steps actually begin in. Keep ids stable — the desktop shell field-sniffs
|
|
# them and tests pin them.
|
|
STEPS: "dict[str, str]" = {
|
|
"env_prefs": "Restoring settings…",
|
|
"native_preload": "Preparing GPU libraries…",
|
|
"ml_imports": "Loading ML runtime (PyTorch)…",
|
|
"api_routes": "Loading API routes…",
|
|
"db_migrate": "Preparing database…",
|
|
"services_start": "Starting background services…",
|
|
}
|
|
|
|
_lock = threading.Lock()
|
|
_t0 = time.monotonic()
|
|
_current: "str | None" = None
|
|
_done: "list[tuple[str, float]]" = [] # (step_id, seconds it took)
|
|
_started_at: float = 0.0
|
|
_ready = False
|
|
_error: "dict | None" = None
|
|
|
|
|
|
def begin_step(step_id: str) -> None:
|
|
global _current, _started_at
|
|
with _lock:
|
|
_finish_current_locked()
|
|
_current = step_id
|
|
_started_at = time.monotonic()
|
|
|
|
|
|
def _finish_current_locked() -> None:
|
|
global _current
|
|
if _current is not None:
|
|
_done.append((_current, round(time.monotonic() - _started_at, 2)))
|
|
_current = None
|
|
|
|
|
|
def mark_ready() -> None:
|
|
global _ready
|
|
with _lock:
|
|
_finish_current_locked()
|
|
_ready = True
|
|
|
|
|
|
def fail(message: str) -> None:
|
|
"""Record a startup failure against the step that was running."""
|
|
global _error
|
|
with _lock:
|
|
_error = {"step": _current, "message": str(message)[:500]}
|
|
|
|
|
|
def is_ready() -> bool:
|
|
with _lock:
|
|
return _ready
|
|
|
|
|
|
def current_step() -> "tuple[str | None, str | None]":
|
|
"""(step_id, human label) of the active step, or (None, None)."""
|
|
with _lock:
|
|
if _current is None:
|
|
return None, None
|
|
return _current, STEPS.get(_current, _current)
|
|
|
|
|
|
def snapshot() -> dict:
|
|
"""The `/startup/progress` body. Always safe to call, never raises."""
|
|
with _lock:
|
|
if _error is not None:
|
|
status = "failed"
|
|
elif _ready:
|
|
status = "ready"
|
|
else:
|
|
status = "starting"
|
|
states = {sid: "pending" for sid in STEPS}
|
|
for sid, _t in _done:
|
|
states[sid] = "done"
|
|
if _current is not None:
|
|
states[_current] = "active"
|
|
if _error is not None and _error.get("step"):
|
|
states[_error["step"]] = "failed"
|
|
durations = dict(_done)
|
|
return {
|
|
"status": status,
|
|
"step": _current,
|
|
"label": STEPS.get(_current, _current) if _current else None,
|
|
"steps": [
|
|
{
|
|
"id": sid,
|
|
"label": label,
|
|
"state": states.get(sid, "pending"),
|
|
**({"t": durations[sid]} if sid in durations else {}),
|
|
}
|
|
for sid, label in STEPS.items()
|
|
],
|
|
"elapsed_s": round(time.monotonic() - _t0, 2),
|
|
"error": _error,
|
|
}
|
|
|
|
|
|
def _reset_for_tests() -> None:
|
|
global _current, _ready, _error, _started_at
|
|
with _lock:
|
|
_current = None
|
|
_done.clear()
|
|
_ready = False
|
|
_error = None
|
|
_started_at = 0.0
|