Commit Graph
1353 Commits
Author SHA1 Message Date
Aleksander Grygier 05cc00a7a0 ui : target the selected model on external backends
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier 26042d1c59 ui : skip llama.cpp-only requests and fields on external backends
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier b0f67285d1 ui : honor per-backend chat and models paths
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier 6aa2c26f18 ui : switch backends from the models selector
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier b6866d47d2 ui : gate llama.cpp-only features by backend capabilities
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier 65a3ed656b ui : stop settings save from clobbering backend config
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier dfcf80f7b3 ui : list models from a backend
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier cfac7cfd13 ui : allow excluding the local backend
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier 808cff003e ui : add backend management settings UI
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier a6b0f69a93 ui : add backend presets, CRUD and connection test
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier 1c36525771 ui : resolve API requests against the active backend
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier fd3b265ffe ui : add backend store and API base resolution
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier 79795f3fba ui : add backend types and constants
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:10:10 +02:00
Aleksander Grygier 6bc13953a8 ui : Improvements
Adds a measured-px height transition to the tool category groups in the
Tools submenu, replacing the bits-ui Collapsible whose conditional
rendering kills the transition. The same technique as the reasoning panel:
the rows stay mounted, the height animates between 0 and the measured
scrollHeight, and visibility keeps collapsed rows out of the tab order.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 12:10:05 +02:00
Pascal 4446688ced ui : resolve bundled icon paths against the app base path
Serve the recommended MCP server favicons and the Hugging Face badge from
the app's base path, not the domain root, so they resolve when the app is
mounted under a path. Also normalizes the safe HTML config indentation.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 12:10:05 +02:00
Aleksander Grygier 7d046fa690 ui : sort reasoning panel attributes
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:09:32 +02:00
Aleksander Grygier 04d7c1eaf9 ui : lighten the model selector rows
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:09:32 +02:00
Aleksander Grygier 8fa7f16e47 ui : cache hugging face hub model data per session
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 12:09:32 +02:00
Aleksander Grygier df1b9d95b2 ui : group the model selector components
Move the four selector surfaces into models/ModelsSelector with their own
barrel, extract the shared reasoning panel and download row, group the option
list helpers under navigation/utils, and add the models selector hook wiring
(in-flight downloads and reasoning menu included).

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 12:09:32 +02:00
Aleksander Grygier 5f2d7ed21c ui : models discover dialog
Add the Models Discover explorer behind a new sidebar action and full-screen
dialog: a searchable HuggingFace GGUF list (org avatars, capability icons,
catalog size ranges) and a detail pane with header, metadata chips, the
sanitized model-card readme and the download area - per-quant action chips
grouped by bit depth, sidecar queuing and a scrollable serve-command preview
with dynamic quant selects. Searchable dropdowns gain sticky headers/footers.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 11:40:29 +02:00
Aleksander Grygier 516afbb9b7 ui : render shared model row hints as native titles
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 11:40:29 +02:00
Aleksander Grygier 3870f15875 ui : remember hub avatars that failed to load
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 11:40:29 +02:00
Aleksander Grygier 7f1c9236a4 ui : shared model display primitives
Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 11:40:29 +02:00
Aleksander Grygier f0127f0c0c ui : optional sanitized raw HTML in markdown
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 11:40:28 +02:00
Aleksander Grygier 85ecd8ce35 ui : model download pipeline
Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 11:40:28 +02:00
Aleksander Grygier 31fc7a0f76 ui : model memory-fit estimation
Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 11:40:28 +02:00
Aleksander Grygier 61ccf5b821 ui : strip provider tilde prefix from hub avatar urls
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 11:40:28 +02:00
Aleksander Grygier 8e37618aaf ui : Hugging Face Hub data layer
Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 11:40:28 +02:00
Aleksander Grygier 774b215cd7 ui : model id grammar for sidecars, quants and capability parsing
Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 11:40:28 +02:00
Aleksander Grygier b5f7cbeb1c ui : type-safe API types, fetch helpers and download-ready models store plumbing
Assisted-by: pi:GLM-5.3-Flash
2026-09-28 11:40:28 +02:00
Aleksander Grygier d1cfcee644 server : fix deadlock when removing a finished download
The download monitor thread acquires the mutex on its way out, so joining
it while holding the lock in server_models::remove deadlocks once the
status has flipped to DOWNLOADED. Join outside the lock, same pattern as
load_models().

Assisted-by: pi:zai-org/GLM-5.3
2026-09-28 11:40:28 +02:00
Sarah Wu 0c6a6a7ce5 Enables Windows ARM64 build with MSVC cl.exe (#28362)
* can reproduce the issue vlad sees

* fix fma issue

* drop volatile

* fix volatile runtime task

* add arm flag if needed

* fix hsum compile error

* fix syntax in quants

* strengthen sve probing

* make the syntax fixes one liners

* remove debug code

* formatting

* remove macro for float

* drive down gcc instruction count

* support armec

* fix CI comments

address CI comments

fix cross compile issue

remove warning

fix style and fix fma probing

fix style

* add documentation

* update documentation
2026-09-28 10:07:27 +02:00
Tim Wangandtimothywang21 4da6337767 server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)
* server : allow splitting RANK pooling for causal LLM rerankers

Rerank models fall into two categories: bidirectional cross-encoders
(BERT, etc.) that require all tokens in a single physical batch, and
causal LLMs repurposed as rerankers (Qwen3, Qwen3-VL) that can use
chunked prefill like any other decoder.

Previously the server rejected all RANK-pooling inputs larger than
n_ubatch, and the graph builder hardcoded QWEN3/QWEN3VL arch checks to
determine last-token pooling. This broke long-document and multimodal
reranking for causal models.

Fix: expose llama_get_causal_attn(ctx) so the server can check the
effective runtime attention type (reflecting any --attention override
or set_causal_attn call). Also expose llama_model_is_causal(model)
for querying the static architectural property from GGUF metadata.

can_split() now permits chunked prefill for RANK pooling when the
context is causal. The graph builder's inline arch check is replaced
with the same cparams.causal_attn predicate, removing the duplication.

Assisted-by: Opencode/Qwen3.8-27B

* remove unused llama_model_is_causal, fix whitespace

Assisted-by: opencode

---------

Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
2026-09-27 23:28:10 +02:00
Georgi Gerganov a97cce86a8 common : avoid side effects around params parsing (#29537)
- register --rpc unconditionally and call llama_supports_rpc() only from its handler
- print server "initialization ..." log after args are parsed

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
2026-09-27 20:18:56 +03:00
Adrien Gallouët 187664b537 llama-bench : fix OOB access of hf_file (#29515)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-27 10:24:54 +02:00
Adrien Gallouët fcb3074f2b server : fix wake_fd warning on Windows (#29479)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-26 21:25:25 +02:00
Vladislavandplotnikov.v10 965f89794f polished Readme and llama-bench (#28968)
Co-authored-by: plotnikov.v10 <plotnikov.v10@wb.ru>
2026-09-26 14:21:21 +08:00
Adrien Gallouët 4b1a27fa0e common,rpc : simplify fs_create_directory_with_parents() (#29432)
The original function was broken on Windows for some unicode paths

Paths without a trailing separator now create the last directory too,
matching the function name. All current callers already include a
trailing separator, so this change does not affect them.

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-25 20:33:37 +02:00
Yuri Khrustalev fcc891545b mtmd: fix mel preprocessor in LFM2 audio (#29403)
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese
test utterances. In Japanese, some differences changed entire words.

This change:

* uses `log(x + 2^-24)` instead of clamping to the log floor
* uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)`
* adds the normalization epsilon to the standard deviation instead of inside the square root

Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged.

Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against
http://github.com/Liquid4All/liquid-audio fp32.

Test set:

* 200 LibriSpeech `test-clean` utterances (EN)
* 200 Common Voice `ja` test utterances (JP)
* identical 16 kHz audio passed to both implementations

| Greedy transcript identical to `liquid-audio` | Without fix |    With fix |
| --------------------------------------------- | ----------: | ----------: |
| EN F16                                        |     191/200 | 200/200 |
| JP F32                                        |     187/200 | 200/200 |
| JP F16                                        |     187/200 | 199/200 |

The remaining JP F16 difference is a comma and matches the reference implementation's own bf16
output.

Mel relative L2 error versus `liquid-audio`:

* EN: 3.2% -> ~2e-6 median
* JP: 3.9% -> ~2e-6 median
2026-09-25 18:52:44 +02:00
Adrien Gallouët 27b20ba8b1 common : extract shared unicode path/string helpers (#29415)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-25 11:40:41 +02:00
Nandan Vallamdasu 945064fcea ui : fix missing svg use and animation elements in preview and download (#28962)
* ui : allow svg use and animation tags in sanitizer

* ui : neutralize href animation retargeting in svg sanitizer
2026-09-24 17:26:01 +02:00
Daniel Bevenius 308883b335 server : change default pytest workers to 4 (#29376)
This commit changes the default number of pytest workers to 4 instead of
auto.

Refs: https://github.com/ggml-org/llama.cpp/pull/29369#issuecomment-5815050948
2026-09-24 15:59:13 +02:00
Tobyandaetherbird f830688e91 model : add Ling 3.0 VL support (#29151)
* model : fold Ling 3.0 VL into the BailingMoeV3 architecture

Assisted-by: Scout

* model : keep shared NORM rope list intact when gating bailingmoe3 on mrope sections

---------

Co-authored-by: aetherbird <aetherbird@users.noreply.github.com>
2026-09-24 08:57:31 +02:00
Adrien Gallouët 2b70583997 server,common : fix the GCC 12 stringop-overread false positive (again) (#29325)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-24 08:40:06 +02:00
Xuan-Son Nguyen b9ae43a5d4 server: allow preset to set log file (#29334) 2026-09-24 01:16:00 +02:00
Will 42916d83f4 server: fix token counting API crash on sleep (#29309)
* server: wake up sleeping server correctly

* server: wake up sleeping server correctly (local aliases removed)
2026-09-23 15:28:49 +02:00
YiChen Lv ee3ecce05c metal : key the fa-vec tuned table by family instead of SKU (#29075)
* key the fa-vec tuned table by family instead of SKU

* fall back to baseline for untuned fa-vec gpu families
2026-09-23 19:23:15 +08:00
calebrio02 057494f93f server: accept OpenAI video_url content type and data: video URIs (#27921)
The OpenAI chat completions API specifies content part type "video_url"
with a {"url": ...} object, and clients typically send data: URIs
(e.g. data:video/mp4;base64,...). The llama-server only accepted the
non-standard "input_video" type and rejected data: URIs for video
(accept_base64_uri=false), so any OpenAI-conformant client failed with
"unsupported content[].type" or "Invalid uri format".

- accept "video_url" as an alias of "input_video"
- read the media object from whichever key was used
- allow data: URIs for video (data:video/*), as already done for images
2026-09-23 12:58:09 +02:00
Xie Wenxiang bcbc936a87 server: Dedup the draft HF model via dedup-cache-models (#27934)
* server: Dedup the draft HF model via dedup-cache-models
Fixes #27846

* server: avoid capturing structured binding in lambda
2026-09-23 12:57:53 +02:00
Pascal 9919911185 server: fix router eviction races with the existing queue (#29217)
* server: route every model load through the queue

A model loaded by the fast path has no queue entry, so tick() evicts
it at its LOADED transition before its own request is proxied. Every
load now joins the queue, whose entry protects the model until its
waiters leave.

* server: do not admit requests into a stopping model

A request for a model that is being stopped still sees it LOADED and
is proxied into the dying child. Such a request now joins the queue
and is served by the next instance. The stopping mark is cleared
under the same lock that sets UNLOADED, so no request can see a
model that is neither stopping nor unloaded while its child is gone.
2026-09-22 21:38:54 +02:00