* feat(downloads): Xet fast path + accurate progress (FDL W0–W2)
Make model downloads fast and show accurate downloaded/remaining/speed.
Research confirmed hf-xet already implements the IDM/uGet technique
(content-defined chunking, parallel byte-range gets, dedup, resume), and
the spike found all 25 catalog repos are Xet-backed — so the win is
driving Xet well + accurate progress, not a custom downloader.
W1 — maximize + guarantee Xet:
- pin huggingface_hub>=1.7 + hf-xet>=1.1 (was transitive); no hf_transfer
- drive snapshot_download with explicit tqdm_class + max_workers + endpoint
- opt-in HF_XET_HIGH_PERFORMANCE / HDD sequential-write knobs (default off)
- /system/info reports fast_download {xet_enabled, xet_version, high_perf}
W2 — accurate progress:
- dry_run preflight -> install_plan event (exact total/cached/remaining)
- utils/download_aggregator.py: one overall bar; byte bars (by id) vs the
"Fetching N files" count bar; windowed rate; emits one 'aggregate' event
- frontend overall bar (speed/remaining/ETA), cached-skip, ⚡ fast badge
Known limit (verified live): under Xet+hf_hub 1.7.2 per-file byte bars
never advance/close via tqdm, so mid-download the bar is file-granular and
bytes flush to the exact total on completion. Classic-LFS/mirror repos get
true byte progress (W4).
Drive-by: download.py used os.walk without importing os (latent NameError
in _validate_snapshot_has_weights on every install) — fixed.
Tests: tests/backend/setup/test_download_preflight.py (10). Spike + plan
under .planning/quick/260613-fdl-fast-model-downloads/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in mirror + cancel + docs (FDL W4)
- mirror (FDL-10): snapshot_download(endpoint=) honours prefs hf_endpoint /
env HF_ENDPOINT on preflight + download (per-call, no process-wide env).
Documented as the classic-LFS path (no Xet) for restricted networks.
- cancel (FDL-11): POST /models/install/cancel {repo_id} stops further
retries at the next boundary, emits install_cancelled, clears the cooldown
(cancel is intent, not failure). Frontend treats it as a terminator.
- docs (FDL-12): docs/downloading-models.md (Xet fast path, progress
semantics + byte-speed limitation, opt-in tuning, mirror, cancel,
troubleshooting) + README pointer. Docs-sync rule satisfied.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(planning): model-management v2 cleanup plan (mm2)
GSD plan for cleaning the model-management subsystem: registry unload-on-
switch + per-engine unload() (fixes VRAM leak), model_lifecycle facade,
unified idle/timeout config, bounded cooldowns, sidecar VRAM self-report,
cache-fallback logging. Planning artifact only — no code.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(downloads): reconcile with main's HF_HUB_DISABLE_XET; honest status
Rebasing onto main surfaced that main forces HF_HUB_DISABLE_XET=1 (classic
LFS) because Xet progress bypasses the tqdm hook — the same limitation found
here. Reconcile instead of fight:
- /system/info fast_download now reports runtime truth: xet_installed +
xet_active (installed AND not HF_HUB_DISABLE_XET) + xet_enabled alias. The
⚡ badge only shows when Xet actually runs; startup log says
"downloads: Xet disabled → legacy LFS".
- complete(): clear the rate window before the final flush so crediting the
full size in one step can't emit an absurd instantaneous rate.
- docs/downloading-models.md rewritten: default is legacy LFS for accurate
progress; Xet is opt-in via HF_HUB_DISABLE_XET=0. hf-xet pin stays (ready
for a future Xet progress hook).
W2 (preflight total/remaining + aggregate bar + exact completion) is the
value on either path; W1's "maximize Xet" is dormant by main's design.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(downloads): opt-in segmented multi-connection accelerator (FDL W3)
Since main forces Xet off (HF_HUB_DISABLE_XET=1), the default path is
single-stream legacy LFS — so a segmented downloader is the way to get BOTH
parallel speed and live byte progress.
- services/segmented_download.py: async multi-connection Range downloader for
one file — parallel byte-ranges, resume (.part + manifest), per-segment
short-read truncation guard, optional sha256/etag verify, cancel, and a
single-stream fallback when the server won't range. Auth-safe: the HF
Authorization header is sent only to huggingface.co/hf.co and never
forwarded to a CDN host on redirect (unit-tested).
- dispatch (download.py): opt-in via prefs segmented_downloader / env
OMNIVOICE_SEGMENTED_DOWNLOAD (default off). When on and Xet inactive,
fetches each file into the HF cache mirroring hf_hub_download (blobs +
snapshot symlinks + refs/main), feeding real bytes to the aggregator. Any
failure falls back to snapshot_download — never breaks a correct install.
- fix: complete() was adding a full total on top of accumulated segmented
bytes (2x); now replaces byte bars so the sum is exactly total.
Verified live (accelerator on): real byte progress to ~16.6 MB/s, final
bytes==total, /models installed=True, delete frees correctly.
Tests: test_segmented_download.py (7) + aggregator double-count regression.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(downloads): relocate FDL tests to top-level; loop-isolate segmented test
CI runs the full suite, which exposed a pre-existing test-isolation leak:
several tests/backend/** fixtures purge core.*/services.* from sys.modules
under a temp OMNIVOICE_DATA_DIR and never restore, leaving core.config/core.db
bound to a dead temp dir. It only bites when collection order puts a purging
test ahead of a real-DB reader (test_longform_jobs). Adding tests under
tests/backend/setup/ reordered collection and tripped it.
Fix without touching the shared (fragile) fixtures or risking class-identity
breakage from a blanket sys.modules restore:
- move the two FDL test files to top-level tests/ (tests/test_fdl_*.py) so
tests/backend/** collection order is identical to main — longform passes.
- rewrite the segmented test to run each case under asyncio.run() (fresh loop)
instead of asyncio.get_event_loop(), which an earlier async test can leave
closed in the full suite.
Full suite green locally: 1364 passed, 0 failed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: mergetest <test@local>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
225 lines
8.5 KiB
Python
225 lines
8.5 KiB
Python
"""Opt-in multi-connection segmented downloader (FDL-08).
|
|
|
|
An IDM/uGet-style downloader for a single file: it fetches many byte-ranges in
|
|
parallel over HTTP, resumes a partial download, verifies the result, and can be
|
|
cancelled. It exists for the **legacy-LFS** download path — the app forces that
|
|
path by default (``HF_HUB_DISABLE_XET=1``) because Xet's progress is opaque, and
|
|
classic LFS is single-stream, so this restores parallel speed *and* keeps live
|
|
byte progress (it reports every received chunk to the aggregator).
|
|
|
|
Auth safety (critical): the Hugging Face ``Authorization`` header is sent **only**
|
|
to ``huggingface.co``/``hf.co`` hosts. When a resolve URL redirects to a CDN
|
|
(CloudFront/etc.), the presigned URL already carries auth, so the token is
|
|
**never** forwarded to the CDN host. Redirects are followed manually to enforce
|
|
this per-hop.
|
|
|
|
This module is deliberately framework-free and unit-tested with
|
|
``httpx.MockTransport``; the HF-cache integration lives in the setup router.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
import json
|
|
import os
|
|
from typing import Callable, Optional
|
|
|
|
import httpx
|
|
|
|
# Hosts the HF token may be sent to. Anything else (CDN) gets no auth header.
|
|
_HF_AUTH_HOSTS = ("huggingface.co", "hf.co")
|
|
_DEFAULT_CONNECTIONS = 8
|
|
_MIN_SEGMENT_BYTES = 4 * 1024 * 1024 # don't split below this — overhead > gain
|
|
_READ_CHUNK = 1024 * 1024
|
|
|
|
|
|
class DownloadCancelled(Exception):
|
|
"""Raised when ``cancel_check()`` returns True mid-download."""
|
|
|
|
|
|
def _host_gets_auth(url: str) -> bool:
|
|
host = (httpx.URL(url).host or "").lower()
|
|
return host in _HF_AUTH_HOSTS or host.endswith(".huggingface.co")
|
|
|
|
|
|
def _auth_headers(url: str, token: Optional[str]) -> dict:
|
|
if token and _host_gets_auth(url):
|
|
return {"Authorization": f"Bearer {token}"}
|
|
return {}
|
|
|
|
|
|
async def _resolve(client: httpx.AsyncClient, url: str, token: Optional[str], max_hops: int = 10):
|
|
"""Follow redirects manually, dropping the auth header on any cross-host hop.
|
|
|
|
Returns (final_url, size_or_None, accepts_ranges_bool).
|
|
"""
|
|
cur = url
|
|
for _ in range(max_hops):
|
|
r = await client.head(cur, headers=_auth_headers(cur, token))
|
|
if r.status_code in (301, 302, 303, 307, 308) and "location" in r.headers:
|
|
cur = str(httpx.URL(cur).join(r.headers["location"]))
|
|
continue
|
|
r.raise_for_status()
|
|
size = r.headers.get("content-length")
|
|
size = int(size) if size and size.isdigit() else None
|
|
accepts = r.headers.get("accept-ranges", "").lower() == "bytes"
|
|
return cur, size, accepts
|
|
raise httpx.TooManyRedirects(f"exceeded {max_hops} redirects for {url}")
|
|
|
|
|
|
def _plan_segments(size: int, num_connections: int) -> list[tuple[int, int]]:
|
|
n = max(1, min(num_connections, max(1, size // _MIN_SEGMENT_BYTES)))
|
|
step = -(-size // n) # ceil
|
|
segs = []
|
|
start = 0
|
|
while start < size:
|
|
end = min(start + step, size) - 1
|
|
segs.append((start, end))
|
|
start = end + 1
|
|
return segs
|
|
|
|
|
|
def _manifest_path(part: str) -> str:
|
|
return part + ".done"
|
|
|
|
|
|
def _load_done(part: str, size: int) -> set[tuple[int, int]]:
|
|
try:
|
|
with open(_manifest_path(part)) as f:
|
|
data = json.load(f)
|
|
if data.get("size") != size:
|
|
return set()
|
|
return {tuple(s) for s in data.get("done", [])}
|
|
except (OSError, ValueError):
|
|
return set()
|
|
|
|
|
|
def _save_done(part: str, size: int, done: set) -> None:
|
|
try:
|
|
tmp = _manifest_path(part) + ".tmp"
|
|
with open(tmp, "w") as f:
|
|
json.dump({"size": size, "done": sorted(list(s) for s in done)}, f)
|
|
os.replace(tmp, _manifest_path(part))
|
|
except OSError:
|
|
pass
|
|
|
|
|
|
async def segmented_download(
|
|
url: str,
|
|
dest: str,
|
|
*,
|
|
token: Optional[str] = None,
|
|
expected_size: Optional[int] = None,
|
|
expected_etag: Optional[str] = None,
|
|
num_connections: int = _DEFAULT_CONNECTIONS,
|
|
on_bytes: Optional[Callable[[int], None]] = None,
|
|
cancel_check: Optional[Callable[[], bool]] = None,
|
|
client: Optional[httpx.AsyncClient] = None,
|
|
timeout: float = 30.0,
|
|
) -> str:
|
|
"""Download ``url`` to ``dest`` using parallel byte-ranges with resume.
|
|
|
|
- ``on_bytes(delta)`` is called as bytes land (feeds the aggregator).
|
|
- ``cancel_check()`` is polled between chunks; returning True raises
|
|
:class:`DownloadCancelled` and leaves the ``.part`` for a later resume.
|
|
- On success the size (and ``expected_etag`` if given) is verified, then the
|
|
``.part`` is atomically renamed to ``dest``.
|
|
"""
|
|
own_client = client is None
|
|
client = client or httpx.AsyncClient(follow_redirects=False, timeout=timeout)
|
|
part = dest + ".part"
|
|
|
|
def _cancelled() -> bool:
|
|
return bool(cancel_check and cancel_check())
|
|
|
|
try:
|
|
final_url, probed_size, accepts_ranges = await _resolve(client, url, token)
|
|
size = expected_size or probed_size
|
|
|
|
# Single-stream fallback: server won't range, or we don't know the size.
|
|
if not accepts_ranges or not size:
|
|
await _stream_single(client, final_url, token, part, on_bytes, _cancelled)
|
|
else:
|
|
done = _load_done(part, size)
|
|
_preallocate(part, size)
|
|
segments = [s for s in _plan_segments(size, num_connections) if s not in done]
|
|
lock = asyncio.Lock()
|
|
|
|
async def _fetch(seg: tuple[int, int]):
|
|
start, end = seg
|
|
want = end - start + 1
|
|
headers = {**_auth_headers(final_url, token), "Range": f"bytes={start}-{end}"}
|
|
async with client.stream("GET", final_url, headers=headers) as r:
|
|
r.raise_for_status()
|
|
got = 0
|
|
with open(part, "r+b") as fh:
|
|
fh.seek(start)
|
|
async for chunk in r.aiter_bytes(_READ_CHUNK):
|
|
if _cancelled():
|
|
raise DownloadCancelled()
|
|
fh.write(chunk)
|
|
got += len(chunk)
|
|
if on_bytes:
|
|
on_bytes(len(chunk))
|
|
# Truncation guard: preallocation makes the file `size` bytes
|
|
# regardless of what arrived, so the per-segment received count
|
|
# — not the file size — is what proves the bytes are real.
|
|
if got != want:
|
|
raise ValueError(f"segment {start}-{end} short read: got {got}, want {want}")
|
|
async with lock:
|
|
done.add(seg)
|
|
_save_done(part, size, done)
|
|
|
|
if segments:
|
|
await asyncio.gather(*(_fetch(s) for s in segments))
|
|
|
|
# ── verify ──────────────────────────────────────────────────────
|
|
actual = os.path.getsize(part)
|
|
if size and actual != size:
|
|
raise ValueError(f"size mismatch: got {actual}, expected {size}")
|
|
# etag is typically the sha256 (LFS) — verify when it looks like a hash
|
|
if expected_etag:
|
|
tag = expected_etag.strip('"')
|
|
if len(tag) == 64 and all(c in "0123456789abcdef" for c in tag.lower()):
|
|
if _sha256(part) != tag.lower():
|
|
raise ValueError("sha256 mismatch — download corrupt")
|
|
|
|
os.makedirs(os.path.dirname(dest) or ".", exist_ok=True)
|
|
os.replace(part, dest)
|
|
try:
|
|
os.remove(_manifest_path(part))
|
|
except OSError:
|
|
pass
|
|
return dest
|
|
finally:
|
|
if own_client:
|
|
await client.aclose()
|
|
|
|
|
|
async def _stream_single(client, url, token, part, on_bytes, cancelled) -> None:
|
|
async with client.stream("GET", url, headers=_auth_headers(url, token)) as r:
|
|
r.raise_for_status()
|
|
with open(part, "wb") as fh:
|
|
async for chunk in r.aiter_bytes(_READ_CHUNK):
|
|
if cancelled():
|
|
raise DownloadCancelled()
|
|
fh.write(chunk)
|
|
if on_bytes:
|
|
on_bytes(len(chunk))
|
|
|
|
|
|
def _preallocate(part: str, size: int) -> None:
|
|
# Create/extend the file to `size` so segment writes can seek to offsets.
|
|
with open(part, "a+b") as fh:
|
|
fh.seek(0, os.SEEK_END)
|
|
if fh.tell() < size:
|
|
fh.truncate(size)
|
|
|
|
|
|
def _sha256(path: str) -> str:
|
|
import hashlib
|
|
h = hashlib.sha256()
|
|
with open(path, "rb") as f:
|
|
for block in iter(lambda: f.read(1024 * 1024), b""):
|
|
h.update(block)
|
|
return h.hexdigest()
|