d91beef0fd314250d8d9b94de86dfea019a8bd96
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
19ae20111a |
fix(security): replace persistent admin keys with scoped sessions (#1528)
* fix(security): replace persistent admin keys with sessions Exchange the remote administrator key once for bounded, revocable credentials. Canonicalize backend principals, enforce cookie CSRF and exact origins, and use path-bound one-use WebSocket tickets. Migrate the bundled UI away from durable master-key storage and credential-bearing URLs. Add unit, integration, static-hygiene, and production-browser regressions plus synchronized operator documentation. * docs: link session hardening to PR 1528 * fix(security): key session indexes with process pepper Use HMAC-SHA-256 instead of an unkeyed digest for in-memory session and WebSocket-ticket indexes. This preserves constant-size lookup identifiers, makes copied records unusable without the process pepper, and resolves CodeQL's weak sensitive-data hash finding. * fix(auth): align empty bearer migration precedence Centralize the Authorization-channel presence decision with canonical principal parsing. Bearer followed only by spaces now remains an empty channel during legacy cookie migration, while unsupported or invalid explicit credentials stay authoritative and fail closed. * fix(security): harden admin session review boundaries * fix(security): derive key generations with HKDF * fix(auth): anchor the admin-session store so module reloads cannot fork it test_master_exchange_does_not_bypass_pin_on_normal_routes failed in full-suite runs: test_mcp_bindings' client fixture purges the services.* tree from sys.modules and reloads main, so api.routers.auth re-imported a fresh services.admin_sessions (new AdminSessionStore) while core.auth kept its import-time reference to the old one — the exchange issued the cookie into one store and the middleware resolved it against another, turning the expected "PIN required" into "API key required". Root cause is the class of bug, not the one test: a process-global auth store defined as a bare module-level singleton forks under importlib.reload or purge-and-reimport. Fix at the source: admin_session_store now resolves through a synthetic sys.modules anchor (_omnivoice_admin_session_store_anchor) that reloads never re-execute and package-prefix purges never match, so every copy of the module shares the one per-process store. No consumer or behavior changes. Regression test reproduces both fork vectors (in-place reload and sys.modules purge + fresh import) and asserts previously issued sessions still resolve and the store identity is preserved; it fails before this fix and passes after. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(auth): honor X-Forwarded-Proto for CSRF origin and Secure cookies behind TLS proxies Behind Tailscale Serve (docs/remote-gpu.md) or any TLS-terminating proxy, the browser talks https while the backend hop stays http, so exact-origin CSRF compared an https Origin against an http expectation and rejected every legitimate request, and the session cookie shipped without Secure. uvicorn's ProxyHeadersMiddleware only rewrites the scope for loopback peers, which misses Docker and any non-loopback proxy topology. New core.csrf.effective_scheme derives the client-facing scheme: resolved scope first (uvicorn's trusted-proxy rewrite wins), then an upgrade-only read of X-Forwarded-Proto's first value — https/wss promotes http to https, everything else is ignored, and a genuine TLS hop can never be downgraded. Used by both the destination-origin comparison and auth._secure_cookie so the WS-ticket/logout CSRF paths and the cookie Secure flag agree. Spoofing gains nothing: the host:port half of the origin tuple is untouched, browsers cannot attach the header cross-site without a preflight this API never grants, and forging it on plain http only adds Secure (the browser then drops the cookie — self-harm only). Regression tests: proxied https origin accepted (origin check, Secure flag, logout), comma-separated chains, scope-fallback path, spoofed header still rejects cross-origin, cannot downgrade real https, junk values ignored. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(auth): consume the stored admin key only after a successful exchange A remote-backend user upgrading with their backend unreachable lost the only stored copy of OMNIVOICE_API_KEY: every migration path deleted the durable ov_api_key BEFORE the session exchange settled, stranding them until they recovered the key from the server box. Close the whole class: - client.ts bootstrap: read the legacy key, exchange first, and remove the durable copy only after the exchange succeeds; on failure the key stays so the next launch retries the migration (auth gate still rises). - authSession.ts exchangeApiKey: move removeLegacyMaster from before the fetch to the cookie/bearer success paths — the key never coexists with a live session, but a rejected or hung exchange no longer consumes it. - remoteBackendProbe.ts configuredRemoteBackend: stop wiping the key on every app mount. - RemoteBackendPanel: a connection test or an aborted save no longer wipes the pending key; only disabling the remote backend discards it. - prefKeys.js: ov_api_key moves from PREF_KEYS to PRESERVED_KEYS — factory reset preserves the pending connection credential exactly like ov_backend_url; the successful migration is what deletes it. Tighten the credential-hygiene static guard to match: it accepted sessionStorage.setItem('ov_api_key', …) — the exact class it exists to close. The guard now flags .setItem(<master key>) on any storage receiver, quote style, or injected-store alias, with a self-test pinning what it catches and what stays legal. Fail-before/pass-after regression tests: backend unreachable retains the key and the next bootstrap retries it; a successful exchange removes it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * perf(auth): make session validation occupancy-independent * test(auth): catch optional master-key storage calls * feat(docs): add PR control document for bultodepapas in VoiceStudio * docs: keep the PR tracking board in the fork; credit the changelog line The pr-control document is excellent process discipline, but it is the contributor's own operational board (their inventory, their update commands) — it lives naturally in their fork, and docs/agents/ here is context every repo agent loads. Removed with appreciation; the changelog line gains its contributor credit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: debpalash <4178343+debpalash@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d49c56df4a |
fix(shell): root error boundary so an app-render throw can't blank the window
The mount chain (StrictMode → QueryClientProvider → RemoteAuthGate → App) had no error boundary above App. App's per-tab <ErrorBoundary> wraps only cover their own subtrees, so a throw in App's own render — its top-level hooks, store access, restoreProjectExtras, or the header/sidebar chrome that renders before any tab — escaped every boundary and left #root empty. The shell's blank_guard could then only reload three times and paint a dead-end failure page: exactly the "#root children = 0 after 3 reloads" a user hit. Wrap the mount root in <ErrorBoundary name="app-root"> so such a throw shows the in-app, recoverable error card (Reload / Report, data untouched) instead of a blank window. The shell guard stays as the last resort for the rarer case where even the boundary can't render. Recurrence-proofing (the throw only had to happen once to blank the app, and the #1178-class variant only bites the MINIFIED bundle that dev + jsdom never run): - main-app.test.jsx: fail-before/pass-after — a throwing app tree must leave #root non-empty (the recovery card), not blank. - Wire the existing production-bundle smoke (e2e-prod/prod-bundle-smoke.spec.ts) into CI so a pre-render crash in the shipped bytes fails the pipeline, not a user's launch. Fixed its IGNORABLE list to treat the no-backend WebSocket handshake error as the harness artifact it is. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ecfe42569c |
style: oxfmt the prod-bundle gate files
format:check gates CI and the two new files were not oxfmt-clean. Scoped to just those files: running `bun run format` wholesale on Windows rewrites all 479 files' line endings, which would bury the change in CRLF churn. |
||
|
|
939b4d4b78 |
test+dev: production-bundle black-screen gate, and stop dev instances colliding
Two independent causes of a blank app window. Never SHIP a blank (the prod hole). v0.3.22 shipped a black screen: a minifier temporal-dead-zone reorder threw before React mounted, leaving an empty #root (#1178). It reached users because every existing check — vitest, node:test, and the whole Playwright e2e suite — runs the UN-MINIFIED dev server, so a bug living only in the minified bundle passes them all. playwright.prod.config.ts + e2e-prod/ close that hole: build the real bundle, serve dist/ via vite preview, and assert the app actually mounts (#root has children, renders visible text, no pageerror). The core assertion is deliberately structural — "did anything mount?" — because that is what a pre-render crash always breaks, whatever its cause. retries: 0, so a blank screen can never be flaky-passed away. Never DISPLAY a blank in dev (the collision). Running `bun desktop` while one is already up does not fail politely, it cascades into a blank window. Reproduced deterministically and measured over CDP: the healthy main window has #root childElementCount 1; after a second launch it is 0. The new launch's port grab makes the running instance's dev:api exit, and `concurrently --kill-others-on-fail` then tears down that instance's whole stack including its Vite server — leaving its window open, pointed at a dev URL that no longer answers. desktop-dev.mjs now clears a leftover dev app first, loudly. The safety boundary for that cleanup is `isDevAppProcess` in desktop-common.mjs: it matches the cargo dev binary (`omnivoice-studio`) ONLY, never the installed release app (`OmniVoice Studio`) — killing a user's real app would be far worse than the bug being fixed. Unit-tested both ways. Also makes the gate runnable off Linux: the dev e2e config hardcodes /usr/bin/chromium, which doesn't exist on Windows/macOS. The new config falls back to Playwright's own browser so a contributor can run the gate before a release. The ci.yml step that runs this gate is NOT in this commit — pushing workflow changes needs a token scope this session lacks. It is provided separately for the maintainer to apply. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |