Commit Graph
1294 Commits
Author SHA1 Message Date
Aleksander Grygier cdac3a729d feat: WIP 2026-09-04 20:07:36 +02:00
Aleksander Grygier 10f985a047 feat: WIP 2026-09-04 20:07:36 +02:00
Aleksander Grygier 065e8d883a chore: Add LLAMA-APP-REUSE comments 2026-09-04 20:07:36 +02:00
Aleksander Grygier c024c1c567 feat: Models Downloading UI/UX 2026-09-04 20:07:36 +02:00
Aleksander Grygier bd1f21efb3 ui : show downloaded quants and drafts as static chips
Downloaded files are no longer selectable toggles; they render as a
chip with a checkmark instead.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 5ef15ee2ce ui : merge download options and terminal commands into one panel
The quant and draft sidecar buttons become a multiple toggle group
(selecting which files to download, not download triggers), and below
the panel the terminal llama serve command updates live to reflect the
selection, with a download CTA that fires the downloads for the
selected entries (main + optional draft sidecar).

Adds the shadcn-svelte toggle-group component.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 6be54d7bd6 ui : add download manager pieces and the list search component
- ModelsDiscoverListSearch: extracted search input from the list
- ModelsDiscoverModelDetailsMetadata: description + metadata chips
  extracted from the details header
- ModelsDiscoverModelDetailsCommands: quant + draft sidecar selectors
  embedded in the inline command text
- ModelsDownloadManager: tracked downloads with per-file progress and
  a delete action
- ModelsDownloadManagerDownloadStatusToast: one toast per download
  with a progress bar per file (main + sidecars) and a CTA to open
  the download manager
- DialogModelsDownloadManager: dialog shell for the manager

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier e436b5e96c ui : restructure discover components to the target names
- ModelsDiscoverItem + ModelsDiscoverInfo fold into
  ModelsDiscoverListItem (avatar + model id + badges + context/size)
- ModelsDiscoverDetails* renamed to ModelsDiscoverModelDetails*
- TerminalCommands renamed to ModelsDiscoverModelDetailsCommands
- ModelsDiscoverDetailsName folded into the details header
- ModelsDiscoverListSearch extracted from the list search input
- stories updated for the new names

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier bd2a47e27f ui : add a Discover models trigger to the model selector
Add a Discover models item to the selector dropdown footer that opens
the DialogModelsDiscover dialog.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier b5d4ab37c5 ui : wire the discover dialog to live data and downloads
Load the selected model's details, file tree and README via
HuggingFaceService on selection change, and expose the download
progress type globally for the status feed. The download options
already read their state from the models status store.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 09e47cd45a ui : simplify download options to plain memory estimate
Replace the device-memory tier badges with the simple memory estimate:
each quant tooltip shows the estimated runtime memory and the device/OS
chip is dropped, matching the estimateModelMemoryBytes model. Make the
download dialog callbacks optional so the component stays presentational
until wired to the live status store.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 94e20a6059 ui : add models discover components with stories
Port the discover UI from the scrapbook, adapted to the typed
sidecar API: searchable two-pane explorer (list, item, info, org
avatar with quant badge), model details (header, name badges,
download options grouped by bit depth with compatibility tiers,
terminal serve/cli commands per draft sidecar, README viewer, chat
template dialog), download confirmation dialog with progress, and
the full-screen dialog shell.

Presentational components take data and download state via props;
the loading container and store wiring land in the integration
branch. MarkdownContent gains a sanitized allowHtml option used by
the README viewer; ModelId gains context, size-range, params and
sidecar badges; the selector option passes thinking/tool flags
instead of the removed capabilities prop.

Basic Storybook stories cover each component with HF-shaped
fixtures.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 17731e9a92 ui : mark llama-app-reusable code with a LLAMA-APP-REUSE tag
Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier d4482116dd ui : mark llama-app-reusable code with a LLAMA-APP-REUSE tag
Tag the pure-logic files and functions that llama.app (llama-pages)
can reuse as-is: model id parsing, HF name and quant conventions,
hardware compatibility estimation, chat-template capability
detectors, and the HF formatting and metadata helpers. App-specific
code is left unmarked.

The LLAMA-APP-REUSE prefix makes the reusable surface greppable and
distinguishable from regular comments: grep -rn LLAMA-APP-REUSE.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier bac2c2d9eb ui : fix double blank line in model types
Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 6629357df4 ui : add download state tracking to the model status manager
Port the download lifecycle into ModelsStatusManager: track per-entry
progress keyed by <repo>:<tag> from the /models/sse feed, record
failed downloads for the delete-and-retry path, and expose the
downloadModel / cancelDownload operations (POST/DELETE /models).

Add the ServerModelStatus.DOWNLOADED/DOWNLOADING cases and the
ModelDownloadProgress type.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier fb3e6277d1 ui : add model download pipeline
Wire the model download flow: ModelsService.downloadModel (POST
/models) and cancelDownload (DELETE /models), the apiDelete helper,
ApiModelsDownloadRequest/Response types, the download_progress SSE
payload, and the download_finished/download_failed SSE event kinds
matching the server feed.

Add modelsHubStore owning the HuggingFace GGUF model list for the
discover dialog: curated catalog defaults on open, search replaces
the list across all of HuggingFace.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier cd24a67c4a ui : simplify model memory estimation
Replace the device-memory tier machinery with a plain file-size
estimate: required runtime memory is the model file size with
headroom for KV cache and allocator overhead (estimateModelMemoryBytes).
Callers present the requirement; there is no device detection and no
fit-versus-budget verdict.

Drops resolveDeviceMemoryGb, deviceMemoryBudgetMb,
computeFileCompatibilityTiers and the CompatibilityTier type, and the
barrel keeps only the new estimator.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier bd90ba5d7a ui : add model compatibility estimation
Port the hardware-compatibility estimator from ggml-org/llama-macos:
map every GGUF file in a repo to a full/limited/none tier based on
the device memory budget (GPU working set approximated from RAM, less
fit slack and an OS floor) and the estimated weight + context memory.

Main quants are tiered individually; shards, mmproj and quant-matched
draft sidecars inherit their main quant's tier. Sidecar picking
mirrors the server's find_best_sibling ranking (deepest directory,
exact quant tag, closest bit depth).

Also port detectToolUseSupport (infers tool-calling support from a
chat template) and the browser get_info fallback helper.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 41d3821e0d ui : move catalog url to models-discover constants
Address review follow-up: the llama.app catalog endpoint belongs to
the models-discover feature, not the HF constants. Use Number() for
shard index parsing and name the UD-quant prefix segment lookup.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier da017bcb53 ui : move HF constants to constants and enums modules
Address review on the HF data layer:

- replace the HfModelSort / SidecarForm / sibling entry type string
  unions with enums (HfModelSort, SidecarForm, HfEntryType)
- move URLs, query params, regexes, limits, retry settings, shard
  file conventions, tag tokens and formatting units into a dedicated
  huggingface.constants.ts; reuse the existing PATH_SEPARATOR
- drop the task label / pipeline icon / library display maps: the
  discover UI only presents GGUF models, so keep the task tags for
  logic use only (parseTags)
- drop the hardcoded curated model list; the discover dialog gets its
  default list from the llama.app /v1/catalog.json endpoint, which is
  an acceptable online-only source since the feature requires internet
  access anyway

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 8d496326f7 ui : follow MODEL_ID regex rename in huggingface service
Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier 0871357fe6 ui : add huggingface hub data layer
Add HuggingFaceService for browsing and searching GGUF models on the
HF Hub: catalog/model search, model details, repo file tree, raw
README fetch, and the llama.app model catalog. Includes GGUF file
analysis helpers - extractQuantMeta (quant token plus sidecar type
and its form, prefix or suffix), shard collapsing, quant bit-depth
lookup, and download/size/likes formatting.

Add the HF API types and the curated model list shown in the
Discover Models sidebar.

Assisted-by: pi
2026-09-04 20:07:36 +02:00
Aleksander Grygier ad8e42f71a ui : lowercase reasoning capability and use sidecar enums in tests
Assisted-by: pi
2026-09-04 20:07:35 +02:00
Aleksander Grygier bb44a904a4 ui : address review on model id parsing
Use lowercase values for the sidecar enums so the value doubles as
the filename token, derive the sidecar regexes from the enum values,
and rename the MODEL_ID regex keys to the _REGEX suffix used by the
rest of the constants files. Replace the tools capability magic
string with ModelCapability.TOOL_USE.

Assisted-by: pi
2026-09-04 20:07:35 +02:00
Aleksander Grygier 2bf5c2c837 ui : parse sidecar types in model ids
Add ModelDraftSidecar / ModelAuxSidecar enums with a ModelSidecar
union type; mmproj is the only auxiliary sidecar (single member,
covers vision and audio input). Add SIDECAR_PREFIX/SUFFIX_RE regex
matching the server's filename conventions, and type guards +
enum-file-token helpers in model-id.constants.ts.

Extend parseModelId to detect sidecar filename tokens (mtp-, mmproj-,
etc) and expose isDraftSidecar / isAuxSidecar / sidecarFromFileToken
helpers. Add ModelCapability.TOOL_USE with icon/label/flag mappings.

Assisted-by: pi
2026-09-04 20:07:35 +02:00
Aleksander Grygier 693f870034 ui : add format:files script for formatting changed files only
Assisted-by: pi
2026-09-04 20:07:35 +02:00
Aleksander Grygier 01da32baf2 ui : rename model API types to mirror endpoints
Drop the Router prefix from client-side API types; names now map
directly to the /models endpoint family (load/unload/download/list).
Merge ApiModelListResponse into ApiModelsListResponse (same endpoint
shape in both modes) and remove the duplicate ModelsService.listRouter().

Assisted-by: pi
2026-09-04 20:07:35 +02:00
Aleksander Grygier 2d64ea1290 ui : remove dead ChatFormActionAddMcpSubmenu component
Unreferenced leftover from the pre-dialog MCP design; references
context props that no longer exist and breaks svelte-check.

Assisted-by: pi
2026-09-04 20:07:35 +02:00
nachobh 85d5703a3b ui : fix MCP image attachments not displayed in tool block (#25789) (#28089)
* ui : fix MCP image attachments not displayed in tool block (#25789)

Fixes regression from #25450 where ChatMessageAgenticContent passed
message.extra instead of section.toolResultExtras to tool blocks,
leaving tool images invisible. Also fixes TOOL_RESULT_JSON_OPEN_REGEX
which misclassified "[Attachment saved: ...]" as JSON.

Fixes #25789

Assisted-by: Muse Spark

* Addressed PR comments: 1.- Removed ·?? mesage?extra· as it has no case left to cover 2.- Added ·[\· to cover the case of ·[[1, 2], [3, 4]]· case suggested in the PR comment 3.- Added unit test for covering up this regex case

* ui : fix MCP image attachments not displayed in tool block (ggml-org#25789) - Addressed lint error on regex (redundant \)
2026-09-04 19:53:16 +02:00
Tom Tan 1863ac0333 ui: export conversations from database instead of cached store (#27432) 2026-09-04 15:13:10 +02:00
Xuan-Son Nguyen 163a40796f model, mtmd: fix gemma4 vision handling (#28335)
* model, mtmd: fix gemma4 vision handling

* nits
2026-09-04 12:23:27 +02:00
Evan Huus d509cb1e86 Don't use npx inside a package.json script (#28270) 2026-09-04 10:27:56 +02:00
Daniel Bevenius 42f0225fea server : use pytest-xdist for server tests (#28298)
* server : use pytest-xdist for server tests

This commit adds pytest-xdist to the server tests. This is pytest
plugin that distributes test execution across multiple CPU cores.

Assisted-by: pi:llama.cpp/qwen3.8-27B

Refs: https://github.com/ggml-org/llama.cpp/pull/26734#issuecomment-5220707042

* remove server_base_port and BASE_PORT

* use worksteal and pytest builting tmp_path
2026-09-03 15:04:30 +02:00
Xuan-Son Nguyen de8656bd94 mtmd: propagate const to preproc class (#28310) 2026-09-03 12:57:10 +02:00
Mads Marquart f45576aa86 mtmd : add const in various places (#28307)
* mtmd : mark context as const in more methods

Mark `mtmd_context` as `const` in:
- mtmd_bitmap_init_lazy
- mtmd_tokenize
- mtmd_tokenize_from_parts
- mtmd_helper_support_video
- mtmd_helper_bitmap_init_from_file
- mtmd_helper_bitmap_init_from_buf
- mtmd_helper_video_init
- mtmd_helper_video_init_from_buf
- mtmd_helper_model_can_chat

The tokenization functions in particular are useful to have marked
`const`, as that allows more easily telling the compiler that we can
safely tokenize from multiple threads (`mtmd_tokenize` is already
documented as thread-safe, this just reifies that in the signature).

* mtmd : mark tokenization input pointer as const

Mark the `bitmaps` and `parts` pointers in `mtmd_tokenize` and
`mtmd_tokenize_from_parts` as `const`. This allows more easily calling
these with immutable arrays / vectors.

* mtmd : mark llama_context as const in mtmd_helper_model_can_chat
2026-09-03 12:12:49 +02:00
Xuan-Son Nguyen 67a17c17ca mtmd: fix idefics3 preproc (#28273) 2026-09-03 01:00:57 +02:00
Abhiram 9cffdcc801 server : accept data: URLs for input_video and input_audio (#27735)
* server : accept data: URLs for input_video and input_audio

input_video and input_audio passed accept_base64_uri=false to
handle_media(), so data: URLs got treated as raw base64 strings and
failed later with a confusing media probe error (#27724).

pass true for these two content types the same way image_url already
does, and allow video/audio mime types in the data: url check instead
of image only. data URL validation now throws std::invalid_argument so
malformed input comes back as 400 instead of 500, matching the other
input validation in this file.

* server : simplify handle_media and drop unused accept_base64_uri flag

* server : update comment and add unit test for invalid data URI MIME
2026-09-02 22:24:31 +02:00
Xuan-Son Nguyen 7339054744 mtmd: add mtmd_tokenize_from_parts() (#28250)
* add mtmd_tokenize_from_parts

* use it in mtmd-cli

* move add_special to call level
2026-09-02 21:20:10 +02:00
Xuan-Son Nguyen 9400c8946e model: correctly support input vision for deepseek4 (#28154)
* model: correctly support input vision for deepseek4

* nits
2026-09-02 19:14:46 +02:00
Georgi GerganovandXuan-Son Nguyen e750b887a8 common, server : enable preserve_reasoning kwarg by default, log its effective state (#28174)
* common, server : enable preserve_reasoning kwarg by default, log its effective state

If the preserve_reasoning chat template kwarg is not specified explicitly
via --reasoning-preserve / --no-reasoning-preserve, it is enabled by
default after argument processing. The server logs the effective state of
the kwarg, warns that it is enabled by default when the template supports
it, and only warns "has no effect" when it was enabled explicitly on a
template that does not support it. Setting the kwarg via
--chat-template-kwargs is deprecated.

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* cont : update comment

Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>

---------

Co-authored-by: Xuan-Son Nguyen <son@huggingface.co>
2026-09-02 19:19:54 +03:00
Xuan-Son Nguyen 7798007a29 mtmd: support DeepSeek-V4-Flash-Vision-Exp (#28133)
* mtmd: support DeepSeek-V4-Flash-Vision-Exp

* handle min/max token counts from CLI

* rm debugging

* use GGML_ROPE_TYPE_VISION

* nits

* apply review comments

* correct token count
2026-09-02 16:43:43 +02:00
Pascal 0f3a71be15 mtmd: Fix Qwen3-tts-0.6b (#28231)
* mtmd: load the qwen3-tts code predictor proj_in as optional

The talker and the code predictor share the hidden size on the 0.6B
checkpoints, so the reference builds no small_to_mtp_projection and
the conversion emits no tensor for it. The graph already falls back
to identity when the weight is missing, the loader now agrees.

* mtmd: keep the qwen3-tts code predictor ffn_down in F32

The code predictor carries a massive activation: its layer 2 FFN
intermediate peaks around 1.5e5, well past the 65504 ceiling of F16.
mul_mat casts its input to the weight type, so an F16 ffn_down turns
that peak into inf, the residual follows, and the next rms_norm yields
NaN. Reference forward in float32 gives 145109 against 145396 measured
in the graph.
2026-09-02 12:46:16 +02:00
Pascal 774ee0e200 ui: copy the displayed text of grouped agentic responses (#27832)
* ui: copy the displayed text of grouped agentic responses

Agentic sessions render as a single entry anchored on the first
assistant turn, whose content is typically just the first tool call,
so the copy button wrote an empty string to the clipboard. Derive the
text sections of the whole session and copy them joined, matching the
visible response. Plain messages keep the previous behavior.

* const
2026-08-31 17:48:43 +02:00
hmirinandGeorgi Gerganov a7cc83bbae rpc: avoid serializing buffers from other servers (#26500)
* rpc: avoid serializing buffers from other servers

Only include remote buffer pointers when the buffer belongs to the RPC dispatcher receiving the graph. Add a two-server regression test for cross-server tensor serialization.

Assisted-by: Codex

* cont : add ref

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-08-30 20:26:16 +03:00
Georgi Gerganov bebc9350ec common: rename --tensor-read-lazy to --lazy-mode, add -lzm shorthand (#27969)
Rename the --tensor-read-lazy CLI argument to --lazy-mode, to match the
internal lazy_mode parameter, and add a -lzm shorthand. Sync the READMEs.

Assisted-by: pi:llama.cpp/Qwen3.8-27B
2026-08-30 09:18:10 +03:00
Xuan-Son Nguyen 50f068ffff bench: add --tensor-read-lazy (#27881)
* bench: add --tensor-read-lazy

* rm the alias

* rename to LLAMA_LAZY_MODE_*
2026-08-28 20:51:05 +02:00
BartowskiandXuan Son Nguyen 18443257a3 server: add ctx-per-slot (--kv-unified-per-slot) (#24124)
* Add ctx-per-slot argument for unifid KV cache

* Swap out ctx fractions for ctx pool slots

* Formatting cleanup

* Remove ctx-pool-slots, make ctx-per-slot an int

* refactor it

---------

Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
2026-08-27 22:39:14 +02:00
Xuan-Son Nguyen 732707dff2 quantize: cap working memory size to avoid loading big tensors onto RAM (#27795) 2026-08-27 18:31:13 +02:00
Xuan-Son Nguyen fac889fb38 llama: model_loader: add TENSOR_READ_LAZY (#27794)
* llama: model_loader: add TENSOR_GET_ROW_LAZY

* add --tensor-read-lazy

* rename to TENSOR_READ_LAZY

* gen docs

* address comments
2026-08-27 15:14:34 +02:00