Commit Graph
11039 Commits
Author SHA1 Message Date
Aleksander Grygier ee41ebc15d ui : account for cache tokens on external backends
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 69870519b7 ui : route chat requests through protocol adapters
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 946b08a99c ui : report token stats for external backends
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier a4d6bf249f ui : aggregate models across backends
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier ad967c6165 ui : load backend model lists in the background
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 23a19a7952 ui : only show the server error when no backend is available
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 088030d384 ui : prefetch backend model lists on first selector open
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 5beb29294b ui : send the model on external backend requests
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier e5eaad780f ui : ignore favorites that belong to another backend
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 4e9de83333 ui : target the selected model on external backends
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 2cb944149b ui : skip llama.cpp-only requests and fields on external backends
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 0a390a5b27 ui : honor per-backend chat and models paths
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 9dd4acb71d ui : switch backends from the models selector
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier a68b59e853 ui : gate llama.cpp-only features by backend capabilities
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 5d08b6d65a ui : stop settings save from clobbering backend config
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 12ca72a4d9 ui : list models from a backend
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier a0b5102d91 ui : allow excluding the local backend
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier ca02ab35a1 ui : add backend management settings UI
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 8d6bc71ebd ui : add backend presets, CRUD and connection test
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 0ce4a9cef9 ui : resolve API requests against the active backend
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 31031a0cd2 ui : add backend store and API base resolution
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 1d2cc0daae ui : add backend types and constants
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:22:47 +02:00
Aleksander Grygier 21aecafc8c ui : sort reasoning panel attributes
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:18:32 +02:00
Aleksander Grygier e0a8649e52 ui : lighten the model selector rows
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:08:40 +02:00
Aleksander Grygier 8c2a222edb ui : cache hugging face hub model data per session
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:07:37 +02:00
Aleksander Grygier 445ec8f13d ui : group the model selector components
Move the four selector surfaces into models/ModelsSelector with their own
barrel, extract the shared reasoning panel and download row, group the option
list helpers under navigation/utils, and add the models selector hook wiring
(in-flight downloads and reasoning menu included).

Assisted-by: pi:GLM-5.3-Flash
2026-09-16 19:07:37 +02:00
Aleksander Grygier 375a5b08e9 ui : models discover dialog
Add the Models Discover explorer behind a new sidebar action and full-screen
dialog: a searchable HuggingFace GGUF list (org avatars, capability icons,
catalog size ranges) and a detail pane with header, metadata chips, the
sanitized model-card readme and the download area - per-quant action chips
grouped by bit depth, sidecar queuing and a scrollable serve-command preview
with dynamic quant selects. Searchable dropdowns gain sticky headers/footers.

Assisted-by: pi:GLM-5.3-Flash
2026-09-16 19:07:37 +02:00
Aleksander Grygier 27a11a0f6b ui : render shared model row hints as native titles
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 19:07:15 +02:00
Aleksander Grygier 3712caa69f ui : remember hub avatars that failed to load
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 13:53:34 +02:00
Aleksander Grygier ba804dd783 ui : shared model display primitives
Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.

Assisted-by: pi:GLM-5.3-Flash
2026-09-16 13:53:34 +02:00
Aleksander Grygier 6a4dc7da8e ui : optional sanitized raw HTML in markdown
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.

Assisted-by: pi:GLM-5.3-Flash
2026-09-16 13:53:34 +02:00
Aleksander Grygier 0172e83237 ui : model download pipeline
Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.

Assisted-by: pi:GLM-5.3-Flash
2026-09-16 13:53:34 +02:00
Aleksander Grygier 3e43ff1b33 ui : model memory-fit estimation
Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.

Assisted-by: pi:GLM-5.3-Flash
2026-09-16 13:53:33 +02:00
Aleksander Grygier c9462ce786 ui : strip provider tilde prefix from hub avatar urls
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-16 13:44:40 +02:00
Aleksander Grygier f9f41f6480 ui : Hugging Face Hub data layer
Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.

Assisted-by: pi:GLM-5.3-Flash
2026-09-16 13:12:04 +02:00
Aleksander Grygier 0bdbf547f1 ui : model id grammar for sidecars, quants and capability parsing
Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.

Assisted-by: pi:GLM-5.3-Flash
2026-09-16 13:12:04 +02:00
Aleksander Grygier e5278550ff ui : type-safe API types, fetch helpers and download-ready models store plumbing
Assisted-by: pi:GLM-5.3-Flash
2026-09-16 13:12:04 +02:00
Aleksander Grygier 5963c000ef server : fix deadlock when removing a finished download
The download monitor thread acquires the mutex on its way out, so joining
it while holding the lock in server_models::remove deadlocks once the
status has flipped to DOWNLOADED. Join outside the lock, same pattern as
load_models().

Assisted-by: pi:zai-org/GLM-5.3
2026-09-16 13:12:04 +02:00
Aleksander Grygier 0949acee48 common : resolve <quant>-<sidecar> download tags and list cached sidecars
A Q4_0-mtp style tag now resolves the sidecar file when no model file matches it, so a solo draft or mmproj download actually pulls the file. Cached sidecar files list as their own entries so the state survives a restart, and removing such a tag deletes only the sidecar.

Assisted-by: pi:zai-org/GLM-5.3
2026-09-16 13:12:04 +02:00
y198 60199339bc rpc : invalidate cached compute graph when a referenced buffer is freed (#24292)
The server caches the most recent compute graph per device so that
GRAPH_RECOMPUTE can re-execute it without resending tensor data. The
cached graph nodes hold direct pointers to backend buffers that were
live at graph_compute() time. If any of those buffers is later
released via FREE_BUFFER, the next GRAPH_RECOMPUTE re-executes the
cached graph through the dangling pointers (use-after-free).

The bug is reachable by an unauthenticated remote client. The
dangling pointers point into chunks an attacker can reshape via
subsequent ALLOC_BUFFER/SET_TENSOR commands, and the resulting
read/write through the cached graph is sufficient to leak libc
addresses and hijack the buffer iface vtable used by BUFFER_CLEAR,
yielding remote code execution.

Discard all cached graphs in free_buffer(). The existing null-check
in graph_recompute() then rejects the request and the client falls
back to GRAPH_COMPUTE on the next call.

No protocol or API change.
b11000
2026-09-16 14:03:11 +03:00
Gaurav Garg b04d4e567c Change max context length for auto-fitting with unified KV (#28849) b10999 2026-09-16 16:08:50 +05:30
Aman Gupta 37b53fd454 qwen4exp: add hc ops (#28901) b10998 2026-09-16 16:00:01 +08:00
WenqiangJia2026 fccf7166fb HIP: broaden MoE ncols_opt tile heuristic on RDNA3.5 architecture (#28935)
It's found the MoE ncols_opt tile heuristic needs to be broadened
to include the RDNA3.5 architecture.

The code change is implemented in ggml/src/ggml-cuda/mmq.cu
and just change the GGML_CUDA_CC_IS_RDNA3_0 to
GGML_CUDA_CC_IS_RDNA3 in the condition.
The dense dispatch logic remains unchanged.
The Test machine configuration we used is
AMD Radeon 8060S, gfx1151 (RDNA3.5), 20 CU, wave32
+ AMD Ryzen AI MAX+ 388, 8C/16T, 23.79 GB RAM

we complete the Correctness verification and performance evaluation as follows:
  test-backend-ops test -b ROCm0 -o MUL_MAT    -p type_a=<q4_K|q5_K|q4_0|q5_0>
  test-backend-ops test -b ROCm0 -o MUL_MAT_ID -p type_a=<q4_K|q5_K|q4_0|q5_0>
  all pass: MUL_MAT 64/64, 29/29, 48/48, 14/14;
            MUL_MAT_ID 84/84, 3/3, 74/74, 3/3

Performance result on target machine:
  LFM2.5-8B-A1B-UD-Q4_K_M  (Q4_K MoE)   +16.198%  [+12.704, +19.799]   8/8
  Qwen1.5-MoE-A2.7B-Q2_K   (Q2_K MoE)    +6.189%  [ +5.245,  +7.141]   8/8
  pooled (16 pairs)                     +11.081%  [ +7.972, +14.279]  16/16

Token generation (tg128) is unchanged on the Q4_K MoE model and +2.188%
[+0.905, +3.488] on the Q2_K one.
b10997
2026-09-16 09:55:02 +02:00
Aldehir Rojas 0bec16e388 chat : force \n</think> on reasoning budget end for qwen3-coder (#28869) b10996 2026-09-16 08:47:28 +02:00
SG-Amadeus d4365d9554 vulkan: make MUL_MAT_ID BN/2 tail unconditional (#28923)
Use BN/2 as the default for BNover2 and as the disabled fallback for BNover4, and remove the enable gate from the MUL_MAT_ID BN/2 branch. The BN/4 branch remains gated by enable_smaller_matrices, while the p.N path is unchanged.
b10995
2026-09-16 08:45:44 +02:00
0a8b29a607 metal: fix NaN in mul_mm_id when activations exceed f16 range (#26223)
* test-backend-ops: reproduce MUL_MAT_ID NaN for activations beyond f16

The Metal mul_mm_id path narrows src1 to `half` for the simdgroup MMA
(`S1 = half` in every instantiation; ggml-metal.metal:10582 and :10595,
mirrored at :10643/:10654 in the tensor-ops path). f16 saturates at
65504, so a model whose activations exceed that produces inf, and
`simdgroup_multiply_accumulate` then turns the whole 8x8 accumulator
tile into NaN. The mul_mv_id path used below `ne21_mm_id_min` (32)
carries the same values in f32 and is correct, as is every CPU path.

This was untestable before: `init_mul_mat_id_tensors` initializes
uniform [-1, 1], so no existing case can drive an operand out of f16
range. `test_mul_mat_id` gains an `amax` parameter (default 1.0f,
preserving the historical init exactly) that scales only the f32
activations, leaving the quantized weights in their normal range.

Six cases: n=16 sits below the mul_mv_id -> mul_mm_id switch and is the
control that must stay green; n=32 and n=64 are above it and fail on
Metal today. Two shapes, because this is not model- or size-specific —
q4_K at 128 experts / 4 active / 4096x2048 mirrors a real model, and
q8_0 at 8 experts / 2 active / 512x256 shows the same failure at
minimal size.

Observed on Apple M2 Max, macOS, llama.cpp b10156:
  MUL_MAT_ID(type_a=q8_0,...,n=32,k=256,amax=100000.000000):
    [MUL_MAT_ID] NaN at index 0 (MTL0=nan CPU=583442.375000) FAIL

The real model behind this is Mistral Small 4 (arch mistral4, 128
experts / 4 active), one of whose layers reaches ~1e5 activations: on
Metal every prefill of >=32 tokens returns an entirely NaN vocabulary,
while <32 tokens is correct.

Note kernel_mul_mm (dense) has the identical conversion at :10273 and
:10286 and is expected to fail the same way; it is not covered here.

Found and written by Claude Opus 5 (via Claude Code).

* metal: fix NaN in mul_mm_id when activations exceed f16 range

kernel_mul_mm_id narrows src1 to `half` for the simdgroup MMA operands
(`S1 = half` in every instantiation). f16 saturates at 65504, so a model
whose activations exceed that produces inf on load, and
simdgroup_multiply_accumulate then propagates NaN across the whole 8x8
accumulator tile. The result is an entirely NaN output — not a precision
loss, a total loss. The mul_mv_id path taken below ne21_mm_id_min (32)
keeps the same values in f32 and is correct, as is every CPU path, so
the same model produces correct logits for short inputs and NaN for
long ones.

Fix: rescale src1 by a power of two so it fits, and undo the scale on
the f32 accumulator at the store. A two-stage reduction computes
max(|src1|) and writes the pair (1/scale, scale) into scratch chained
off the destination buffer, in the same style as the existing tpe/ids
id-mapping scratch. The matmul multiplies on load and on store.

This is exact, not approximate, for two reasons: the dot product is
linear, so one tensor-wide factor commutes through the accumulation;
and the factor is a power of two, so both multiplications are exact in
binary floating point. When max(|src1|) already fits — every model that
works today — the factor is exactly 1.0 and the output is bit-identical
to before. Accumulation was already f32 and is unchanged; only the
operand narrowing was ever the problem.

The reduction is two-stage (256 threadgroups into partials, then one
threadgroup folding them) specifically so it stays bandwidth-bound. A
single-threadgroup version was measured first and cost up to +451%
median on prefill — the scan serialized against an otherwise idle GPU.
It is also dispatched only on the mm path, so decode never pays for it.

Measured on Apple M2 Max, `test-backend-ops perf -o MUL_MAT_ID -b MTL0`,
99 cases, versus the same build without this change:

  n=1/4/8   (mul_mv_id, decode)  : -0.8% / -0.8% / -0.4% median (noise)
  n=32      (mul_mm_id, prefill) : +1.73% median
  n=64                           : +1.30% median
  n=128                          : +1.80% median
  n=256                          : +3.98% median
  n=512                          : +3.74% median, +7.20% worst
  overall                        : +1.14% median

Correctness, same machine:
  - the six new test-backend-ops cases go from 4 FAIL / 2 OK to all OK,
    with the n=16 controls (mul_mv_id path) unchanged;
  - `test-backend-ops -b MTL0` full run: 0 failures, no regression;
  - Mistral-Small-4-119B (arch mistral4, 128 experts / 4 active) now
    generates correctly at the default n_ubatch of 512, in both
    UD-IQ3_S and UD-Q4_K_XL quantizations. Before this, every prefill of
    >= 32 tokens returned an all-NaN vocabulary and only n_ubatch <= 31
    (forcing the mul_mv_id path) worked.

Likely fixes #25722 (mistral4 empty output on Metal above ~300 tokens,
FA on and off, generation degenerating to a single control token — the
signature of argmax over an all-NaN distribution). #20668 may be the
same defect attributed to a bad GGUF.

Note kernel_mul_mm (dense) has the identical narrowing at the
corresponding load sites and is expected to fail the same way; it is
left alone here to keep this change reviewable. Also possible, and left
for later: scaling per output column rather than per tensor, which
would preserve more precision when a single token is the hot one.

Found, diagnosed and fixed by Claude Opus 5 (via Claude Code).

* metal : make requested edits

- remove verbose comments
- explain rationale as requested

Generative AI disclosure: Claude made the edits as requested.

* metal : stack mul_mm_id map0 with amax_part

Implement @ggerganov suggestion to stack amax_part + map0. Mean 2.6% faster (worst -0.7%, best -4.1%). Win grows with batch size. Benchmarked on a hot M2 Max after reboot.

Generative AI disclosure:

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* cont : fix var scope

* cont : comment out tests temporarily

Comment out tess to not break CI temporarily

Assisted-by: Claude Fable 5.1

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b10994
2026-09-16 09:37:40 +03:00
Sigbjørn SkjæretandGeorgi Gerganov 583926e3ac ci : add self-hosted webgpu to hf-jobs (#28712)
* add self-hosted vulkan and webgpu to hf-jobs

* try t4-medium

* cont : adjust cpu backend threads

* try t4-small again

* restore cm jobs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b10993
2026-09-16 08:23:58 +02:00
asbelin e13469a323 llama-bench: support --version to print build info (#28971) b10992 2026-09-16 13:39:43 +08:00
Jhen-Jie Hong 930e2fa599 hexagon: add back missing contiguous fast-path and hvx_copy_uu for each run (#28886) b10991 2026-09-15 16:02:06 -07:00
Trivikram Reddy 72b590d65f hex-cpy: use dma if src and dst are contiguous (#28906) b10990 2026-09-15 15:45:28 -07:00