Extract remaining magic values into named constants/enums:
- HF service: reuse PATH_SEPARATOR for the '/' split, new
HF_SHARED_DRAFT_TOKEN and HF_SIZE_STRING_REGEX constants.
- Model status store: reuse MODEL_ID.QUANTIZATION_SEPARATOR for ':' and
PATH_SEPARATOR for the sidecar cache key; new CLI_FLAGS entries for
--model-draft / -md / --mmproj; a new DownloadStopRequest enum replacing
the 'pause' | 'cancel' string union.
- Reuse the shared PATH_SEPARATOR in download-options.utils.
Assisted-by: llama-ui:Qwen3.8-Flash-Next
Move the hardware compatibility budget constants (MIB_BYTES, MB_PER_GB,
QUANT_WEIGHT, RAM_BUDGET_RATIO, RAM_OVERHEAD_MB, MAC_MEM_TIERS) out of
model-compatibility.ts into a new models-scoped constants file, so the
estimation util imports them from $lib/constants.
Assisted-by: llama-ui:Qwen3.8-Flash-Next
Extract the download-options shared symbols out of the component-local
download-options.utils.ts into their lib-scoped folders: QuantOption,
DownloadEntryState, BitDepthRow and SelectableFile into $lib/types; a new
SelectableFileKind enum (replacing the 'main'/'draft'/'aux' string union)
into $lib/enums; SPEC_TYPE, the bit-depth buckets and the draft label into
a new models-discover-download.constants.ts. The util keeps only its
dedicated functions, with named constants for its magic values.
Assisted-by: llama-ui:Qwen3.8-Flash-Next
Move the ModelsDiscover components into ModelsDiscoverList/ and
ModelsDiscoverDetails/ subfolders (with ModelsDiscoverDetailsDownloadOptions/
nested inside), and the model selector components into a ModelsSelector/
subfolder. Rename ModelsDiscoverModelDetails* to ModelsDiscoverDetails* for
consistency, add per-folder index.ts barrels, and keep shared leaves
(ModelId, ModelBadge, ModelLoadHighlight, ModelsDiscoverAvatar,
DownloadProgressBar) at their folder roots.
Cancelling stops the download and discards the partial files, so the
selector item and the download chip both ask for confirmation first.
Assisted-by: pi:zai-org/GLM-5.3
extractQuantMeta now recognizes a -draft tail after the sidecar token
(Model-MTP-draft.gguf), a bare sidecar token filename (imatrix.gguf) and
suffix sidecars whose head carries no quant (Model-imatrix.gguf). imatrix
is an aux sidecar: it stays in the download chips with a badge, is never
a serve-command option, and mmproj alone drives the Vision capability.
Assisted-by: pi:zai-org/GLM-5.3
The download monitor thread acquires the mutex on its way out, so joining
it while holding the lock in server_models::remove deadlocks once the
status has flipped to DOWNLOADED. Join outside the lock, same pattern as
load_models().
Assisted-by: pi:zai-org/GLM-5.3
A "Download in progress" section above the loaded models lists the
in-flight and paused downloads with a live progress bar and the same
pause, resume and cancel actions as the discover quant chips, with the
same avatar and model id presentation as the regular option rows.
Assisted-by: pi
Every quant chip is now an independent download action with its own
lifecycle: download, retry, pause, resume, cancel and delete with
confirmation. The store distinguishes user-stopped downloads over the
/models/sse feed so they settle silently, keeps paused progress
resumable, resolves the server-registered id on delete and refetches
the model list so the selector stays in sync with downloads.
Drops the selection toggle group, the download CTA and the download
manager dialogs and progress toasts. The serve command preview is now
a standalone widget with its own picks and an addable draft segment.
Assisted-by: pi
Hide the ModelId badge and icon wrappers when empty, normalize its indentation, and tighten the reasoning panel and discover detail spacing.
Assisted-by: pi:zai-org/GLM-5.3
Solo sidecar downloads register in /v1/models under the entry tag, so the state checks that first (normalized against the UD- quant prefix) and keeps the --model-draft / --mmproj args of registered models as a fallback. The chips no longer flip to downloaded while a download is in progress.
Assisted-by: pi:zai-org/GLM-5.3
The inline picks now mirror the selection one-way instead of seeding $state, the default quant seeds once when the file list resolves, and the command only renders while something is selected.
Assisted-by: pi:zai-org/GLM-5.3
Move the expand chevron to the right and swap it for ChevronUp when
open. Group the effort label next to the title.
Assisted-by: llama-ui:Qwen3.8-Flash-Next
Replace the submenu-based dropdown with a flat scrollable list that has
a sticky search header and actions footer. Add avatars and parameter
badges to model options. Introduce ModelsSelectorReasoningPanel for
in-place reasoning effort selection. Adjust dropdown sizing and input
blur styling.
Assisted-by: llama-ui:Qwen3.8-Flash-Next
Replaces the native quant selects with the ui/select component using a
new xs trigger size: h-6, mono text, w-fit so the trigger width follows
the selected quant instead of reserving space for the widest option.
The primary-tinted look and the dashed default-preview state carry over
via trigger classes; SELECT_CLASS is no longer needed.
Assisted-by: Claude Sonnet
Reverts the -shared suffix on the chip badge and drops the shared
prefix from the draft quant select options: both MTP variants read as
plain MTP, distinguished only by their file name tooltip and size.
Assisted-by: Claude Sonnet
With iconsOnNewLine, the tool use / reasoning / modality icons now sit
in the second row, in line with context length and disk size, instead of
competing with the name on the identity line. Icons render through a
shared snippet in both positions; single-line usages keep the icons
inline after the name.
Assisted-by: pi:GLM-5.3-Flash
getBaseModels reads cardData.base_model / base_model: tags, which the HF
list endpoint omits by default; search rows therefore fell back to the
repo org avatar. Adding tags to the expand list lets search rows render
the base org avatar as main with the quant org as the corner badge, like
catalog rows.
Assisted-by: pi:GLM-5.3-Flash
The shared marker moves from a 'shared' suffix on the quant name to the
sidecar badge (MTP-SHARED); draft select options prefix the badge text
so both variants of the same quant stay distinguishable in the dropdown,
and the download CTA shows e.g. 'Download Q4_K_M + MTP-SHARED'.
Assisted-by: pi:GLM-5.3-Flash
The parent had been overwritten with a pre-componentization version,
dropping the shared quant toggles, the download button/command
components and the section's frosted surface styling. Restores the
componentized version and merges the manual refinements into the child
components: primary-tinted quant selects, taller CTA with press scale,
command box accent bar with an absolutely positioned copy button.
Assisted-by: pi:GLM-5.3-Flash
Search results now request the fields the rows render (gguf, siblings,
pipeline_tag, ...) via repeated expand params, so a search row shows the
same badges as a catalog row. The catalog path fetches each repo's tree
and derives a per-repo min/max size range (quants plus draft sidecars)
through the models hub store, with search rows falling back to one lazy
tree fetch per repo. Sizes from llama.app catalog size strings parse via
HuggingFaceService.parseSizeBytes when sizeBytes is absent.
Assisted-by: pi:GLM-5.3-Flash
Unsloth ships MTP draft heads in two layouts: shared- files borrow the
embedding/output weights from the target model, others are
self-contained. extractQuantMeta now flags a standalone 'shared'
segment, and the chips, quant selects and download CTA show e.g.
'Q4_K_M shared' next to the plain quant so the two files are
distinguishable.
Assisted-by: pi:GLM-5.3-Flash
Long commands (draft models, long repo ids) no longer wrap to a second
line: the command row scrolls horizontally, tokens never shrink, and the
copy button stays pinned outside the scroll area.
Assisted-by: pi:GLM-5.3-Flash
One label | value metadata chip (model size, context, architecture,
license) now renders through a dedicated component; the chat template
button and gated badge stay inline in the metadata component as they
are different shapes.
Assisted-by: pi:GLM-5.3-Flash
Replaces the 'Loading model...' text with a static skeleton mirroring
the detail layout (header, metadata chips, download options box, readme
lines). List item skeletons now vary widths per row so the loading list
does not render as identical blocks.
Assisted-by: pi:GLM-5.3-Flash
extractQuantMeta parsed the full sibling path, so repo layouts that nest
sidecars in a folder (e.g. MTP/mtp-Model-Q4_0.gguf) failed the sidecar
prefix match and were classified as main weights.
Assisted-by: pi:GLM-5.3-Flash
The quant and draft sidecar buttons become a multiple toggle group
(selecting which files to download, not download triggers), and below
the panel the terminal llama serve command updates live to reflect the
selection, with a download CTA that fires the downloads for the
selected entries (main + optional draft sidecar).
Adds the shadcn-svelte toggle-group component.
Assisted-by: pi
- ModelsDiscoverListSearch: extracted search input from the list
- ModelsDiscoverModelDetailsMetadata: description + metadata chips
extracted from the details header
- ModelsDiscoverModelDetailsCommands: quant + draft sidecar selectors
embedded in the inline command text
- ModelsDownloadManager: tracked downloads with per-file progress and
a delete action
- ModelsDownloadManagerDownloadStatusToast: one toast per download
with a progress bar per file (main + sidecars) and a CTA to open
the download manager
- DialogModelsDownloadManager: dialog shell for the manager
Assisted-by: pi
- ModelsDiscoverItem + ModelsDiscoverInfo fold into
ModelsDiscoverListItem (avatar + model id + badges + context/size)
- ModelsDiscoverDetails* renamed to ModelsDiscoverModelDetails*
- TerminalCommands renamed to ModelsDiscoverModelDetailsCommands
- ModelsDiscoverDetailsName folded into the details header
- ModelsDiscoverListSearch extracted from the list search input
- stories updated for the new names
Assisted-by: pi
Load the selected model's details, file tree and README via
HuggingFaceService on selection change, and expose the download
progress type globally for the status feed. The download options
already read their state from the models status store.
Assisted-by: pi
Replace the device-memory tier badges with the simple memory estimate:
each quant tooltip shows the estimated runtime memory and the device/OS
chip is dropped, matching the estimateModelMemoryBytes model. Make the
download dialog callbacks optional so the component stays presentational
until wired to the live status store.
Assisted-by: pi
Port the discover UI from the scrapbook, adapted to the typed
sidecar API: searchable two-pane explorer (list, item, info, org
avatar with quant badge), model details (header, name badges,
download options grouped by bit depth with compatibility tiers,
terminal serve/cli commands per draft sidecar, README viewer, chat
template dialog), download confirmation dialog with progress, and
the full-screen dialog shell.
Presentational components take data and download state via props;
the loading container and store wiring land in the integration
branch. MarkdownContent gains a sanitized allowHtml option used by
the README viewer; ModelId gains context, size-range, params and
sidecar badges; the selector option passes thinking/tool flags
instead of the removed capabilities prop.
Basic Storybook stories cover each component with HF-shaped
fixtures.
Assisted-by: pi