With iconsOnNewLine, the tool use / reasoning / modality icons now sit
in the second row, in line with context length and disk size, instead of
competing with the name on the identity line. Icons render through a
shared snippet in both positions; single-line usages keep the icons
inline after the name.
Assisted-by: pi:GLM-5.3-Flash
getBaseModels reads cardData.base_model / base_model: tags, which the HF
list endpoint omits by default; search rows therefore fell back to the
repo org avatar. Adding tags to the expand list lets search rows render
the base org avatar as main with the quant org as the corner badge, like
catalog rows.
Assisted-by: pi:GLM-5.3-Flash
The shared marker moves from a 'shared' suffix on the quant name to the
sidecar badge (MTP-SHARED); draft select options prefix the badge text
so both variants of the same quant stay distinguishable in the dropdown,
and the download CTA shows e.g. 'Download Q4_K_M + MTP-SHARED'.
Assisted-by: pi:GLM-5.3-Flash
The parent had been overwritten with a pre-componentization version,
dropping the shared quant toggles, the download button/command
components and the section's frosted surface styling. Restores the
componentized version and merges the manual refinements into the child
components: primary-tinted quant selects, taller CTA with press scale,
command box accent bar with an absolutely positioned copy button.
Assisted-by: pi:GLM-5.3-Flash
Search results now request the fields the rows render (gguf, siblings,
pipeline_tag, ...) via repeated expand params, so a search row shows the
same badges as a catalog row. The catalog path fetches each repo's tree
and derives a per-repo min/max size range (quants plus draft sidecars)
through the models hub store, with search rows falling back to one lazy
tree fetch per repo. Sizes from llama.app catalog size strings parse via
HuggingFaceService.parseSizeBytes when sizeBytes is absent.
Assisted-by: pi:GLM-5.3-Flash
Unsloth ships MTP draft heads in two layouts: shared- files borrow the
embedding/output weights from the target model, others are
self-contained. extractQuantMeta now flags a standalone 'shared'
segment, and the chips, quant selects and download CTA show e.g.
'Q4_K_M shared' next to the plain quant so the two files are
distinguishable.
Assisted-by: pi:GLM-5.3-Flash
Long commands (draft models, long repo ids) no longer wrap to a second
line: the command row scrolls horizontally, tokens never shrink, and the
copy button stays pinned outside the scroll area.
Assisted-by: pi:GLM-5.3-Flash
One label | value metadata chip (model size, context, architecture,
license) now renders through a dedicated component; the chat template
button and gated badge stay inline in the metadata component as they
are different shapes.
Assisted-by: pi:GLM-5.3-Flash
Replaces the 'Loading model...' text with a static skeleton mirroring
the detail layout (header, metadata chips, download options box, readme
lines). List item skeletons now vary widths per row so the loading list
does not render as identical blocks.
Assisted-by: pi:GLM-5.3-Flash
extractQuantMeta parsed the full sibling path, so repo layouts that nest
sidecars in a folder (e.g. MTP/mtp-Model-Q4_0.gguf) failed the sidecar
prefix match and were classified as main weights.
Assisted-by: pi:GLM-5.3-Flash
The quant and draft sidecar buttons become a multiple toggle group
(selecting which files to download, not download triggers), and below
the panel the terminal llama serve command updates live to reflect the
selection, with a download CTA that fires the downloads for the
selected entries (main + optional draft sidecar).
Adds the shadcn-svelte toggle-group component.
Assisted-by: pi
- ModelsDiscoverListSearch: extracted search input from the list
- ModelsDiscoverModelDetailsMetadata: description + metadata chips
extracted from the details header
- ModelsDiscoverModelDetailsCommands: quant + draft sidecar selectors
embedded in the inline command text
- ModelsDownloadManager: tracked downloads with per-file progress and
a delete action
- ModelsDownloadManagerDownloadStatusToast: one toast per download
with a progress bar per file (main + sidecars) and a CTA to open
the download manager
- DialogModelsDownloadManager: dialog shell for the manager
Assisted-by: pi
- ModelsDiscoverItem + ModelsDiscoverInfo fold into
ModelsDiscoverListItem (avatar + model id + badges + context/size)
- ModelsDiscoverDetails* renamed to ModelsDiscoverModelDetails*
- TerminalCommands renamed to ModelsDiscoverModelDetailsCommands
- ModelsDiscoverDetailsName folded into the details header
- ModelsDiscoverListSearch extracted from the list search input
- stories updated for the new names
Assisted-by: pi
Load the selected model's details, file tree and README via
HuggingFaceService on selection change, and expose the download
progress type globally for the status feed. The download options
already read their state from the models status store.
Assisted-by: pi
Replace the device-memory tier badges with the simple memory estimate:
each quant tooltip shows the estimated runtime memory and the device/OS
chip is dropped, matching the estimateModelMemoryBytes model. Make the
download dialog callbacks optional so the component stays presentational
until wired to the live status store.
Assisted-by: pi
Port the discover UI from the scrapbook, adapted to the typed
sidecar API: searchable two-pane explorer (list, item, info, org
avatar with quant badge), model details (header, name badges,
download options grouped by bit depth with compatibility tiers,
terminal serve/cli commands per draft sidecar, README viewer, chat
template dialog), download confirmation dialog with progress, and
the full-screen dialog shell.
Presentational components take data and download state via props;
the loading container and store wiring land in the integration
branch. MarkdownContent gains a sanitized allowHtml option used by
the README viewer; ModelId gains context, size-range, params and
sidecar badges; the selector option passes thinking/tool flags
instead of the removed capabilities prop.
Basic Storybook stories cover each component with HF-shaped
fixtures.
Assisted-by: pi
Tag the pure-logic files and functions that llama.app (llama-pages)
can reuse as-is: model id parsing, HF name and quant conventions,
hardware compatibility estimation, chat-template capability
detectors, and the HF formatting and metadata helpers. App-specific
code is left unmarked.
The LLAMA-APP-REUSE prefix makes the reusable surface greppable and
distinguishable from regular comments: grep -rn LLAMA-APP-REUSE.
Assisted-by: pi
Port the download lifecycle into ModelsStatusManager: track per-entry
progress keyed by <repo>:<tag> from the /models/sse feed, record
failed downloads for the delete-and-retry path, and expose the
downloadModel / cancelDownload operations (POST/DELETE /models).
Add the ServerModelStatus.DOWNLOADED/DOWNLOADING cases and the
ModelDownloadProgress type.
Assisted-by: pi
Wire the model download flow: ModelsService.downloadModel (POST
/models) and cancelDownload (DELETE /models), the apiDelete helper,
ApiModelsDownloadRequest/Response types, the download_progress SSE
payload, and the download_finished/download_failed SSE event kinds
matching the server feed.
Add modelsHubStore owning the HuggingFace GGUF model list for the
discover dialog: curated catalog defaults on open, search replaces
the list across all of HuggingFace.
Assisted-by: pi
Replace the device-memory tier machinery with a plain file-size
estimate: required runtime memory is the model file size with
headroom for KV cache and allocator overhead (estimateModelMemoryBytes).
Callers present the requirement; there is no device detection and no
fit-versus-budget verdict.
Drops resolveDeviceMemoryGb, deviceMemoryBudgetMb,
computeFileCompatibilityTiers and the CompatibilityTier type, and the
barrel keeps only the new estimator.
Assisted-by: pi
Port the hardware-compatibility estimator from ggml-org/llama-macos:
map every GGUF file in a repo to a full/limited/none tier based on
the device memory budget (GPU working set approximated from RAM, less
fit slack and an OS floor) and the estimated weight + context memory.
Main quants are tiered individually; shards, mmproj and quant-matched
draft sidecars inherit their main quant's tier. Sidecar picking
mirrors the server's find_best_sibling ranking (deepest directory,
exact quant tag, closest bit depth).
Also port detectToolUseSupport (infers tool-calling support from a
chat template) and the browser get_info fallback helper.
Assisted-by: pi
Address review follow-up: the llama.app catalog endpoint belongs to
the models-discover feature, not the HF constants. Use Number() for
shard index parsing and name the UD-quant prefix segment lookup.
Assisted-by: pi
Address review on the HF data layer:
- replace the HfModelSort / SidecarForm / sibling entry type string
unions with enums (HfModelSort, SidecarForm, HfEntryType)
- move URLs, query params, regexes, limits, retry settings, shard
file conventions, tag tokens and formatting units into a dedicated
huggingface.constants.ts; reuse the existing PATH_SEPARATOR
- drop the task label / pipeline icon / library display maps: the
discover UI only presents GGUF models, so keep the task tags for
logic use only (parseTags)
- drop the hardcoded curated model list; the discover dialog gets its
default list from the llama.app /v1/catalog.json endpoint, which is
an acceptable online-only source since the feature requires internet
access anyway
Assisted-by: pi
Add HuggingFaceService for browsing and searching GGUF models on the
HF Hub: catalog/model search, model details, repo file tree, raw
README fetch, and the llama.app model catalog. Includes GGUF file
analysis helpers - extractQuantMeta (quant token plus sidecar type
and its form, prefix or suffix), shard collapsing, quant bit-depth
lookup, and download/size/likes formatting.
Add the HF API types and the curated model list shown in the
Discover Models sidebar.
Assisted-by: pi
Use lowercase values for the sidecar enums so the value doubles as
the filename token, derive the sidecar regexes from the enum values,
and rename the MODEL_ID regex keys to the _REGEX suffix used by the
rest of the constants files. Replace the tools capability magic
string with ModelCapability.TOOL_USE.
Assisted-by: pi
Add ModelDraftSidecar / ModelAuxSidecar enums with a ModelSidecar
union type; mmproj is the only auxiliary sidecar (single member,
covers vision and audio input). Add SIDECAR_PREFIX/SUFFIX_RE regex
matching the server's filename conventions, and type guards +
enum-file-token helpers in model-id.constants.ts.
Extend parseModelId to detect sidecar filename tokens (mtp-, mmproj-,
etc) and expose isDraftSidecar / isAuxSidecar / sidecarFromFileToken
helpers. Add ModelCapability.TOOL_USE with icon/label/flag mappings.
Assisted-by: pi
Drop the Router prefix from client-side API types; names now map
directly to the /models endpoint family (load/unload/download/list).
Merge ApiModelListResponse into ApiModelsListResponse (same endpoint
shape in both modes) and remove the duplicate ModelsService.listRouter().
Assisted-by: pi
* ui : fix MCP image attachments not displayed in tool block (#25789)
Fixes regression from #25450 where ChatMessageAgenticContent passed
message.extra instead of section.toolResultExtras to tool blocks,
leaving tool images invisible. Also fixes TOOL_RESULT_JSON_OPEN_REGEX
which misclassified "[Attachment saved: ...]" as JSON.
Fixes#25789
Assisted-by: Muse Spark
* Addressed PR comments: 1.- Removed ·?? mesage?extra· as it has no case left to cover 2.- Added ·[\· to cover the case of ·[[1, 2], [3, 4]]· case suggested in the PR comment 3.- Added unit test for covering up this regex case
* ui : fix MCP image attachments not displayed in tool block (ggml-org#25789) - Addressed lint error on regex (redundant \)
* server : use pytest-xdist for server tests
This commit adds pytest-xdist to the server tests. This is pytest
plugin that distributes test execution across multiple CPU cores.
Assisted-by: pi:llama.cpp/qwen3.8-27B
Refs: https://github.com/ggml-org/llama.cpp/pull/26734#issuecomment-5220707042
* remove server_base_port and BASE_PORT
* use worksteal and pytest builting tmp_path
* mtmd : mark context as const in more methods
Mark `mtmd_context` as `const` in:
- mtmd_bitmap_init_lazy
- mtmd_tokenize
- mtmd_tokenize_from_parts
- mtmd_helper_support_video
- mtmd_helper_bitmap_init_from_file
- mtmd_helper_bitmap_init_from_buf
- mtmd_helper_video_init
- mtmd_helper_video_init_from_buf
- mtmd_helper_model_can_chat
The tokenization functions in particular are useful to have marked
`const`, as that allows more easily telling the compiler that we can
safely tokenize from multiple threads (`mtmd_tokenize` is already
documented as thread-safe, this just reifies that in the signature).
* mtmd : mark tokenization input pointer as const
Mark the `bitmaps` and `parts` pointers in `mtmd_tokenize` and
`mtmd_tokenize_from_parts` as `const`. This allows more easily calling
these with immutable arrays / vectors.
* mtmd : mark llama_context as const in mtmd_helper_model_can_chat
* server : accept data: URLs for input_video and input_audio
input_video and input_audio passed accept_base64_uri=false to
handle_media(), so data: URLs got treated as raw base64 strings and
failed later with a confusing media probe error (#27724).
pass true for these two content types the same way image_url already
does, and allow video/audio mime types in the data: url check instead
of image only. data URL validation now throws std::invalid_argument so
malformed input comes back as 400 instead of 500, matching the other
input validation in this file.
* server : simplify handle_media and drop unused accept_base64_uri flag
* server : update comment and add unit test for invalid data URI MIME