Commit Graph
11377 Commits
Author SHA1 Message Date
Aleksander Grygier 38c2c9cbc2 ui : lead the discover list with downloads
The discover sidebar lists running downloads above the suggested models,
and steps aside while a query is active. A paused download stays in the
manager table, where it is resumed or dropped.

Assisted-by: pi
2026-10-02 11:59:25 +02:00
Aleksander Grygier 1120c39205 ui : resume model downloads from the pause point
A stopped download keeps its partial file instead of deleting it, so
re-posting the tag resumes from disk. The first progress record now
carries the resume offset, the paused snapshot survives into the
resumed entry, and a pause settles as soon as it is requested.

Assisted-by: pi
2026-10-02 11:59:25 +02:00
Aleksander Grygier 06b21f7bba ui : add the Enable Discover Models setting
The flag gates browsing and downloading HuggingFace GGUF models, so it is
declared on the discover layer; the repo-tree lookup in the manager layer
above reads it from here.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 4881f73d31 ui : put the discover view in the models manager
The manager dialog gains its second view: the Discover Models call to
action fades the table out, slides the title with its compass mark, and
fades the discover surface in; the arrow button returns to the table.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 452d055f91 ui : add the discover models components
The Discover surface: the curated-catalog list with its search and
skeletons, the details pane with its readme, metadata, Hub stats and
download options, and the standalone download progress bar that replaces
the plain one. The quant download button sizes its select with a new xs
trigger, so the select gains that size and the muted box look. The store
stays incomplete when a repo fails, so the next mount retries it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 796fbb571f ui : optional sanitized raw HTML in markdown
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.

Assisted-by: pi:GLM-5.3-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 48b43dd924 common : resolve <quant>-<sidecar> download tags and list cached sidecars
A Q4_0-mtp style tag now resolves the sidecar file when no model file matches it, so a solo draft or mmproj download actually pulls the file. Cached sidecar files list as their own entries so the state survives a restart, and removing such a tag deletes only the sidecar.

Assisted-by: pi:zai-org/GLM-5.3
2026-10-02 11:59:24 +02:00
Aleksander Grygier ee5685f8cd ui : refine the model pane and the manager filters
The pane no longer says the server withholds data: it says the metadata
and the chat template appear once the model is loaded, and the template
block grows with its content instead of scrolling inside a fixed box. The
pane ends with a Delete this model from disk action, confirmed by the same
dialog the table row uses and hovering like a destructive dropdown entry.
The filter toggles list only the modalities the models at hand actually
carry, and starting a new chat from the pane closes the dialog.

Assisted-by: pi
2026-10-02 11:59:24 +02:00
Aleksander Grygier 7bc505dab7 ui : list downloads as their own manager section
The models table lists in-flight and paused downloads in a Downloading
section: flat rows like Loaded models, right after them, with a full
width progress bar along the bottom. The status column shows the percent
and swaps in pause or resume on hover; pause, resume and delete live in
the row's action menu. The selector lists running downloads only.

Assisted-by: pi
2026-10-02 11:59:24 +02:00
Aleksander Grygier 7851efc3b2 ui : default to the Hub metadata and hide avatars when it is off
Turn the Hugging Face Hub models-metadata setting on by default, and when
it is off render no org avatar at all instead of the monogram fallback:
ModelOrgAvatar is the choke point, ModelAvatar stops mounting its wrapper,
and the base-model lookups in the selector and the download row are skipped.

Drop the vestigial "Enable Discover Models" setting from this branch: the
discover UI lives on the discover branch, and nothing here referenced it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier 0f18f69694 ui : move mcp servers to the sidebar rail
The add menu keeps to what it attaches to the conversation, so the MCP
servers dialog moves to the sidebar rail and its menu entries and the
form callback they used go with it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier 93732021e7 ui : drop an orphaned doc comment left by a moved method
Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier cfa83f4a1b ui : move reasoning effort beside the model selector
The reasoning level becomes its own control next to the selector, and
the add menu drops its reasoning submenu.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier ca4cc8a541 ui : rework the selector around its providers
The selector keeps working when a provider is down: the trigger shows the
provider's mark and the org avatar, the banner only appears when no
backend is enabled, and the model ids read through their settings. The
model mark becomes MODEL_ICON everywhere.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier b0d0ee6b50 ui : group the model selector components
The selector surfaces move under models/ModelsSelector with their own
barrel: the dropdown and sheet relocate, the list, option and trigger icon
split out, and the shared list helpers join the navigation utils. The
searchable dropdown gains a sticky footer and per-surface class hooks.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier 5fe9a1ad67 ui : add the models manager
A manage models dialog for the local server: a table that folds each repo's
quants into one row, groups them into families, and sorts from its column
headers. Loaded models lead, then favorites, then the local block, then the
hidden one; sections keep their open state in local storage. The toolbar
filters by capability, modality and context, rows act on favorite, delete
and hide, and the manager opens focused on a model from the sidebar or a
download row. The models store gains the recents, hidden and group-open
state, and warms the Hub records for the context and capability columns.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Xuan-Son Nguyen a4cb4c61fd llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (#29818)
* init conversion

* convert: ok

* model loaded

* add server code

* improve conversion script

* support shared prompt prefix

* add docs, imorove UX a bit

* add vision support

* add openjev tiny model for testing

* add dev docs

* support lev & kev

* clean up

* fix lev noul

* fix py lint

* nits docs

* clarify about not supporting date_facts
b11361
2026-10-02 11:56:04 +02:00
Adrien Gallouët 70849ee82c common : remove fs_open_ifstream() by using u8path() (#29841)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-10-02 11:43:45 +02:00
Adrien Gallouët 8d81559fa7 llama : silence unused-result warnings (#29839)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-10-02 11:07:06 +02:00
Adrien Gallouët 6805ae35df llama : use GGML_ABORT instead of throw (#29840)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-10-02 10:44:44 +02:00
lhez a8c9a4e7cc opencl: use sigmoid f16 for bf16 (#29787) 2026-10-02 11:18:51 +03:00
cwriterandcwriter 392ded6546 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (#29186)
* Adding wide-load mmvq for Q8_0 and esimd dmmv for q8_0

Assisted-by: Codex

* remove guard for q8_0

* remove docs

* Simplify by committing to clean code without fallback

* Add feature flag as requested

Assisted-by: Claude Opus 5

---------

Co-authored-by: cwriter <cwriter@localhost>
2026-10-02 11:14:35 +03:00
Jiwoong Song 9e258a6e0a vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (#28531)
Assisted-by: Claude Opus
b11355
2026-10-02 11:13:20 +03:00
Titaniumtown b933289545 sycl: large register file for D=512 FA vec kernels (#29062)
* sycl: large register file for D=512 FA vec kernels

* tests: add 512-wide FA heads to the perf sweep
2026-10-02 11:12:48 +03:00
Łukasz Ślusarczyk c328acc91d sycl : do not use slow oneDNN reference matmul and fattn (#28985)
* sycl : do not use slow oneDNN reference matmul and fattn

* sycl : probe oneDNN matmul once at device init

Assisted-by: Claude Opus 5
2026-10-02 11:11:03 +03:00
Georgi Gerganov 4e2713c162 qwen4exp : optimize mask constructions (#29824)
* qwen4exp : optimize mask constructions

* cont : apply the same change for GLM5-next
b11352
2026-10-02 11:08:50 +03:00
Georgi Gerganov 631109b34d ggml : add alloc_buffer_n to buffer type interface (#23671)
* ggml : add `alloc_buffer_n` to buffer type interface

Add alloc_buffer_n method to ggml_backend_buffer_type_i
interface, with a public API ggml_backend_buft_alloc_buffer_n.

- Default implementation in ggml-backend.cpp handles multi-buffer
  splitting and tensor allocation via ggml_tallocr
- Meta buffer type provides custom implementation that creates
  per-device sub-contexts and delegates to simple buffer types
- ggml_backend_alloc_ctx_tensors_from_buft now collects tensors
  into a list and delegates to the new API
- Remove temporary ggml_backend_meta_alloc_ctx_tensors_from_buft
- Add NULL alloc_buffer_n to all existing buffer type
  interfaces (cpu, metal, openvino, hexagon, webgpu, zdnn, virtgpu, repack)

Assisted-by: llama.cpp:local pi

* cont : fix `cur_buf_size` init after flushing a buffer

* ggml : add TODO tag for shared buffer split logic

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* tests : add alloc_buffer_n coverage

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* cont : fix compile warnings

* tests : add descriptions for alloc_buffer_n tests

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* ggml : address review comments on alloc_buffer_n

- restore GGML_LOG_ERROR on buffer alloc / tensor init failure in the
  default impl (name the failing tensor)
- check the malloc result and drop the _impl indirection in
  ggml_backend_alloc_ctx_tensors_from_buft
- remove comments that restate the code
- fix the TAG_ALLOC_SHARED_BUFFER_SPLIT typo

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* ggml : add get_alloc_size_n to buffer type interface

- Add ggml_backend_buft_get_alloc_size_n public API
- Add optional get_alloc_size_n callback to ggml_backend_buffer_type_i
- Share tensor->buffer planning between alloc_buffer_n default and get_alloc_size_n default
- Replace unchecked realloc with std::vector in alloc_buffer_n default
- Make ggml_backend_alloc_ctx_tensors_from_buft_size use the new API
- Add test-alloc coverage for get_alloc_size_n

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* cont : report malloc failure
b11351
2026-10-02 11:08:08 +03:00
Aaron Teo 254b177307 ci : fix missing zdnn backend check (#29837)
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
2026-10-02 09:10:45 +02:00
Ruben Ortlam fb4b2737a8 vulkan: add logging to pipeline compile issues (#29794) b11349 2026-10-02 08:27:56 +02:00
Kasimir Tanner 207bdab950 pyproject : add linux platform marker to uv torch source (#29177)
In tool.uv.sources, torch was unconditionally pinned to the custom
pytorch CPU index, which lacks macOS Darwin wheels and causes uv sync
to fail on macOS. Add the sys_platform == 'linux' marker to match the
existing Poetry dependencies configuration.

Assisted-by: Antigravity

Resolves: https://github.com/ggml-org/llama.cpp/issues/29176
2026-10-02 07:32:13 +02:00
kurquhar 5fc4f3c8c7 hexagon: install rebuilt HTP skels (#29828)
* hexagon: install rebuilt HTP skels

Assisted-by: OpenCode

* hexagon: fix HTP skel catalog dependencies

Assisted-by: OpenCode
b11347
2026-10-01 19:07:38 -07:00
Aman Gupta 159c651f57 qwen4exp: fix tests (#29819) b11346 2026-10-02 09:12:25 +08:00
Jhen-Jie HongandMax Krasnyansky a868c3e3c5 hexagon: add q2_k and q3_k quant type support (#29717)
* hexagon: add q2_k and q3_k quant type support

* hex-qk: consistent allocation of src1_row_size

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
b11345
2026-10-01 14:12:25 -07:00
Johannes Gäßler ec7630a640 CUDA: fix 2 broken Volta FA cases (#29803) b11344 2026-10-01 21:55:41 +02:00
Johannes Gäßler 78e2964c23 llama: refer to segment documentation [no ci] (#29074) 2026-10-01 21:52:40 +02:00
Adrien Gallouët f1cee9941b common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (#29816)
See https://gcc.gnu.org/bugzilla/show_bug.cgi?id=101510

Close #29759

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11342
2026-10-01 20:15:02 +02:00
Xuan-Son Nguyen 68e79bd8cd skill: note about model-specific CLI arguments + testings (#29808)
* skill: note about adding model-specific CLI arguments

* add testing instructions
2026-10-01 19:53:23 +02:00
Pascal e358d59178 ci: fix Fusion / metal by updating the qwen4exp baseline (#29812) 2026-10-01 19:57:34 +03:00
Georgi Gerganov 81e39ad343 llama : clamp kpool re-pool bound to existing pools (#29805)
* tests : simplify function signature

* llama : clamp kpool re-pool bound to existing pools

The n_tokens/kpool + n_seqs_unq bound on n_new_g overshoots when a batch
fills the whole cache: n_ctx tokens complete exactly n_ctx/kpool pools, so
the +1 pads new_pool_idxs/new_pool_rep one entry past n_pool_real. Graph
reserve only covers n_pool_real entries, so the first full-context decode
builds bigger tensors than reserved and ggml-alloc demands a graph
reallocation (abort under GGML_SCHED_DEBUG_REALLOC=1).

Clamp the bound to n_pool_real: a ubatch can never mark more pools than
the cache holds, and reserve's n_pool_max already covers that.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD

* cont : cap to n_pool_max
b11339
2026-10-01 19:55:34 +03:00
Yiwei ShaoandMax Krasnyansky dcd387a412 hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685)
* hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA

* hex-cpy: various fixes on top of the concat optimizations

Removed CONCAT_DMA_MIN_ROW logic, it was broken with 64-bit DMA.
While it's kinda silly to use DMA for tiny stuff if that tensor gets mapped to an extended buffer the only way to read it is DMA.

Added missing dma_queue_flush() calls.

Added additional guards for conditions we don't support.

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
b11338
2026-10-01 08:38:37 -07:00
Sam Malayek d775ebf363 server: return HTTP 400 for invalid embedding requests (#29060) b11337 2026-10-01 16:52:43 +02:00
Yu Chengye 2b36825cbc convert : write Gemma embedding scale for DFlash drafts (#29802)
* convert : write Gemma embedding scale for DFlash drafts

A DFlash draft shares the target's token embeddings. Gemma scales them by sqrt(hidden_size) in the forward pass, and the draft config does not state that scale, so the converted draft read unscaled embeddings.

Take the scale from the target config when the draft config has none.

Assisted-by: Claude

* convert : check with get_model_architecture for gemma models
2026-10-01 16:34:32 +02:00
42d958167a cuda : route sm70 to the Turing MMVQ nwarps table (#29753)
* cuda : route sm70 to the Turing MMVQ nwarps table

Volta (sm_70) has no MMVQ parameter table of its own and falls through
to GENERIC, which launches K-quant batch-1 decode (ncols_dst == 1) at
nwarps=4. sm_70 shares TURING's tuning: the K-quant vec_dot prefers
nwarps=2 there. Route sm_70 to the existing MMVQ_PARAMETERS_TURING
table in both the device and the host table selector.

Measured on one Tesla V100 32GB PCIe (PG500-216, driver 580.178.04,
CUDA 12.0.140) with Qwen3.8-27B Q4_K_M, tg128, interleaved A/B in 6
ABBA blocks with paired per-block deltas: +1.091 t/s = +3.17 %
(t = +49.0, all six per-block deltas positive); perplexity
bit-identical (6.3697 +/- 0.04066 both builds, wiki.test.raw). The
patched build's K-quant mul_mat_vec_q kernels launch at nwarps=2
(cubin EIATTR_MAX_THREADS) while Q4_0/Q8_0 stay at nwarps=4, and the
same measurement on the September master base gave +3.84 % (t = 85).

The tuning originates from the V100-focused fork anyei/llamacpp-v100
(MIT), commit b912d1b1e, which carries a dedicated
MMVQ_PARAMETERS_VOLTA table; a cubin-level comparison confirmed that
routing sm_70 to the existing TURING table is equivalent for the
K-quant batch-1 path this change affects, so this is the minimal
2-line form. https://github.com/anyei/llamacpp-v100/commit/b912d1b1e

Original-patch-by: anyei <angelyoelroblesmercedes@gmail.com>

* Update ggml/src/ggml-cuda/mmvq.cu

---------

Co-authored-by: tkittich <tkittich@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
b11335
2026-10-01 15:21:26 +02:00
Mike van LammerenandNiklas Wenzel 13b4d7135a metal : release temporary private transfer buffers (#29777)
* metal : release temporary private transfer buffers

Assisted-by: OpenAI Codex

* metal : fix order and formatting

---------

Co-authored-by: Niklas Wenzel <dev@nikwen.de>
b11334
2026-10-01 16:17:23 +03:00
Masashi Yoshimura 4b1622afb7 webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (#29358) b11333 2026-10-01 22:09:07 +09:00
Georgi Gerganov 869034b4bb llama : fix invalid assert in recurrent memory (#29799) b11332 2026-10-01 14:49:03 +03:00
ynankaniandJohannes Gäßler b56f34ab13 CUDA: Handle compute type for NVFP4 on cublass path (#29173)
* CUDA: Handle compute type for NVFP4 on cublass path

Signed-off-by: ynankani <ynankani@nvidia.com>

* Use BF16 compute type for quantized models if HW allows

Signed-off-by: ynankani <ynankani@nvidia.com>

* Set acc prec to bf16 for nvfp4 as it needs atleast bf16 range

Signed-off-by: ynankani <ynankani@nvidia.com>

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* preserve op_params for per-expert matmul

Signed-off-by: ynankani <ynankani@nvidia.com>

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
b11331
2026-10-01 16:53:52 +05:30
Aman GuptaandGeorgi Gerganov c061df1983 Qwen4Exp: add MTP (#29761)
* Qwen4Exp: add MTP

* remove has_state member, check via ctx_bufs being non-empty

* consistent naming + less verbose comments

* cont : clean-up recurrent memory

* cont : clean-up comments

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b11330
2026-10-01 14:13:27 +03:00
Aman Gupta 66e0c17ee1 llama: fix qwen4exp (#29751)
* llama: fix qwen4exp

* qwen4exp: keep kq_mask input the same shape
2026-10-01 14:13:27 +03:00
Oliver Simons 7677678503 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792)
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.

We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)
2026-10-01 13:05:29 +02:00