* feat(network): automatic Hugging Face endpoint selection — probe, pick, remember
Restricted-network first-runs (the #984 class: huggingface.co unreachable,
user dead-ends before discovering the mirror setting) now self-heal by
default, while explicit endpoint choices are never second-guessed.
- New backend/services/endpoint_race.py: parallel HTTPS reachability +
latency probes of huggingface.co and the hf-mirror.com community mirror
(3s timeouts). Probes are the only signal — no geo-IP, no third-party
calls. Reachable beats unreachable; with both reachable the official
endpoint wins unless the mirror is decisively faster (anti-flap
hysteresis). The pick is cached in prefs and re-raced only on first run,
a network-classified download failure, staleness (>7 days), or an
explicit "Test again".
- Manual mode is sacred: HF_ENDPOINT env, an hf_endpoint pref, or any
explicit Settings pick disables auto-switching entirely;
OMNIVOICE_HF_ENDPOINT_MODE=manual is a hard opt-out.
- Wiring: the wizard preflight races endpoints when nothing is configured
(honest copy when the mirror wins; warn-not-block when nothing is
reachable); Model Store installs and the model-cache auto-repair resolve
their per-call endpoint= through the cached decision, and a
network-classified failure re-races once per repo per process and
retries on the new winner (same guard pattern as the cache-recovery
ladder).
- Settings → Models → Hugging Face mirror gains "Auto (recommended)":
shows the current pick, measured latency, last-checked time, and a
"Test again" button (POST /api/settings/hf-mirror/test). Existing
explicit configs surface as the matching manual mode. Panel notes that
hf_hub checksums every download regardless of endpoint.
- Tests: policy/cache/failover matrices in tests/test_endpoint_race.py,
preflight + settings + repair-failover integration with mocked probers,
HFMirrorPanel mode tests, and a suite-wide conftest guard that pins the
probers so no test can hit the real network.
- Docs: downloading-models.md and install/troubleshooting.md describe the
automatic default and both opt-outs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(changelog): Unreleased entry for automatic HF endpoint selection
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(tests): endpoint-probe pin uses an isolated MonkeyPatch and clears the decision cache; dtype guard tolerates stubbed torch
The autouse probe pin requested the shared monkeypatch fixture, hoisting
its setup earlier for every test and reordering teardown against the fp16
guard — which then ran torch.get_default_dtype() on test_torch_compile_gate's
SimpleNamespace stub. The pin now uses its own MonkeyPatch context and also
clears the prefs-cached endpoint decision per test (one test's auto pick
leaked into other tests' preflight labels on CI ordering). The dtype guard
additionally skips non-module torch stubs outright.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(tests): endpoint env vars can no longer leak out of the mirror-settings suite
set_hf_mirror writes os.environ[HF_ENDPOINT] during the test, and
monkeypatch.delenv(raising=False) on an absent var records nothing to
undo — so the write leaked process-wide and flipped later suites'
preflight network checks into the explicit-endpoint branch (the CI-order
failures). Guaranteed save/restore autouse fixture at the source, plus
defensive env shedding in the preflight suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>