Commit Graph
11313 Commits
Author SHA1 Message Date
Aleksander Grygier a027aa20fe ui : mount the configuration pane as a drawer in the manager
A row click slides the pane in as a fixed-width drawer: the toolbar's
calls to action leave first, the pane is laid out before its first open,
and a model switch fades one out and the next in. The model information
dialog folds into the pane's information tab, and a chat started from
the pane closes the manager and focuses the composer.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier e653003019 ui : add the model configuration pane
The pane shows a model's information, load and inference settings side by
side: what its provider reports, the Hub records it falls back to, the
load controls gated by what the serving backend can do, and the
per-model overrides the load form writes. The slider control arrives
with it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier 4a0f556d1c ui : Improvements
Adds a measured-px height transition to the tool category groups in the
Tools submenu, replacing the bits-ui Collapsible whose conditional
rendering kills the transition. The same technique as the reasoning panel:
the rows stay mounted, the height animates between 0 and the measured
scrollHeight, and visibility keeps collapsed rows out of the tab order.

Assisted-by: pi:GLM-5.3-Flash
2026-09-30 15:31:11 +02:00
Pascal 48e043599b ui : resolve bundled icon paths against the app base path
Serve the recommended MCP server favicons and the Hugging Face badge from
the app's base path, not the domain root, so they resolve when the app is
mounted under a path. Also normalizes the safe HTML config indentation.

Assisted-by: pi:GLM-5.3-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier 76be7c1d63 ui : drop an orphaned doc comment left by a moved method
Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier f847272314 ui : move reasoning effort beside the model selector
The reasoning level becomes its own control next to the selector, and
the add menu drops its reasoning submenu.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier 83ac43f54d ui : rework the selector around its providers
The selector keeps working when a provider is down: the trigger shows the
provider's mark and the org avatar, the banner only appears when no
backend is enabled, and the model ids read through their settings. The
model mark becomes MODEL_ICON everywhere.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier 2f524018e4 ui : group the model selector components
The selector surfaces move under models/ModelsSelector with their own
barrel: the dropdown and sheet relocate, the list, option and trigger icon
split out, and the shared list helpers join the navigation utils. The
searchable dropdown gains a sticky footer and per-surface class hooks.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier bf998c5d84 ui : show every provider's models in the manager
The table grows a section per enabled provider, each with its own mark,
error and loading state, and a provider filter with live repo counts.
Rows gain their capability gates back, the draft column with its
use-as-draft action, and the quant badge names the provider on its rows.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier 1fbe94afc0 ui : add the manage providers view
A provider is added from a preset card or by hand: the dialog probes the
endpoint, reads a refusal as a llama.cpp server that wants a key, and
shows the connection test as a status block. Presets carry their official
artwork and a saved backend keeps its branding through the favicon
fallback. The manager dialog gains its third view, and the settings save
stops clobbering keys that other surfaces write.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier 7d843a39c2 ui : route chat through the active provider
Chat goes through protocol adapters, so an OpenAI-compatible endpoint
speaks its own wire format: per-backend paths and headers, the model on
the request, tools kept on the local server, and token counts synthesized
for endpoints that do not stream their own timings. The server store keeps
the local server's props while another provider is active, a conversation
resolves the provider its model belongs to before sending, and the chat
screen never blocks on the local probe when the install has none.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier 74d8206afe ui : add the providers data layer
A provider is a server entry with a base url, an optional key, a protocol
and the paths its API lives at; the local llama.cpp server is the built-in
one. Backends persist in settings and the active one is restored on load.
Requests resolve against the active backend, model ids become
backend-qualified, and every backend's model list is fetched and cached in
the background. The manager's helpers learn to read a model's drafts,
context and the provider that serves it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:11 +02:00
Aleksander Grygier 245e9cf4c1 ui : put the discover view in the models manager
The manager dialog gains its second view: the Discover Models call to
action fades the table out, slides the title with its compass mark, and
fades the discover surface in; the arrow button returns to the table.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:10 +02:00
Aleksander Grygier 5d87573a3b ui : add the discover models components
The Discover surface: the curated-catalog list with its search and
skeletons, the details pane with its readme, metadata, Hub stats and
download options, and the standalone download progress bar that replaces
the plain one. The quant download button sizes its select with a new xs
trigger, so the select gains that size and the muted box look. The store
stays incomplete when a repo fails, so the next mount retries it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:10 +02:00
Aleksander Grygier 95b40f2a05 ui : optional sanitized raw HTML in markdown
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.

Assisted-by: pi:GLM-5.3-Flash
2026-09-30 15:31:10 +02:00
Aleksander Grygier c1a3e445c9 common : resolve <quant>-<sidecar> download tags and list cached sidecars
A Q4_0-mtp style tag now resolves the sidecar file when no model file matches it, so a solo draft or mmproj download actually pulls the file. Cached sidecar files list as their own entries so the state survives a restart, and removing such a tag deletes only the sidecar.

Assisted-by: pi:zai-org/GLM-5.3
2026-09-30 15:31:10 +02:00
Aleksander Grygier ea612b4842 ui : add the models manager
A manage models dialog for the local server: a table that folds each repo's
quants into one row, groups them into families, and sorts from its column
headers. Loaded models lead, then favorites, then the local block, then the
hidden one; sections keep their open state in local storage. The toolbar
filters by capability, modality and context, rows act on favorite, delete
and hide, and the manager opens focused on a model from the sidebar or a
download row. The models store gains the recents, hidden and group-open
state, and warms the Hub records for the context and capability columns.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-30 15:31:10 +02:00
Adrien Gallouët bdeb855b30 ggml-et : remove useless alloca() (#29663)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-30 15:14:45 +02:00
Pascal 3b3d022b82 ci : fix Models Backend Check by shortening the hrm_text fixture (#29744)
The fixture recycles its two blocks over 8 cache slots, so the fp16
error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU
jobs of Models Backend. Two l-cycles keep every branch of the cycle
loop and halve the error.
2026-09-30 15:03:38 +02:00
185103dcf5 llama: llama_prefetch_rows (#29599)
* llama: llama_prefetch_rows

* llama: support row prefetch on Windows

Apply the Windows port contributed by @praneshgo unchanged.

Source: https://github.com/ggml-org/llama.cpp/pull/29599#issuecomment-5887721014

* avoid exposing llama-mmap in model code, route via llama-impl

* add windows check, only prefetch in lazy mode

* cont : clean-up

* cont : fix build

* cont : clarify padding token for gemma4

---------

Co-authored-by: Pranesh Gonegandla <pranesh.iitp@gmail.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-09-30 20:27:18 +08:00
Aman Gupta 2090f60f0b ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (#29675)
* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)

* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops

Assisted-by: Claude Opus 5.5

* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build

* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
2026-09-30 20:24:59 +08:00
Pascal 90c908d06d cpu: accept BF16 in src1 of mul_mat (#28937)
* cpu: accept BF16 in src1 of mul_mat

ggml_conv_1d_dw builds its im2col in F32 when the kernel is BF16, then
calls ggml_mul_mat(im2col, kernel), which puts F32 in src0 and BF16 in
src1. The CPU backend refused that combination, so it was reported as
unsupported on every backend and never compared against anything.

Widen BF16 into the F32 work buffer, next to the existing packing of F32
into vec_dot_type. This is the arithmetic the Metal mat vec kernel
already uses, both operands promoted to float and accumulated in float,
so the two agree exactly rather than approximately.

Cover it with a conv_1d_dw test over F32, F16 and BF16 kernels, plus
three mul_mat cases with BF16 in src1.

* vulkan: reject BF16 in src1 of mul_mat unless src0 is BF16

supports_op only checked the src1 type for non contiguous tensors, so
a contiguous BF16 src1 was accepted and the pipeline lookup asserted.
The only BF16 src1 path is the BF16 x BF16 multiply, every other src0
type now reports the op as unsupported and the scheduler keeps it on
the CPU.

The BF16 kernel case of the conv_1d_dw test needs the f32 x bf16
mat vec variants of the Metal backend, which land separately.
2026-09-30 14:14:45 +02:00
Daniel Bevenius 8df332de1b model-conversion : add --add-bos to run org model script (#29558)
This commit adds an optional --add-bos token command line option to the
run-org-model.py script.

The motivation for this is that there are models, for example Gemma4,
that explicitely set the add_bos value to true in llama-vocab.cpp even
if the original model does not set this value to True.

It would be nice to be able to force the models to agree on the bos
token so that logit verification can proceed.

Refs: https://github.com/ggml-org/llama.cpp/pull/21500
2026-09-30 12:52:02 +02:00
Aleksander Grygier 4a096b8ff6 ui : shared model display primitives (#29644)
* ui : shared model display primitives

Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.

Assisted-by: pi:GLM-5.3-Flash

* ui : remember hub avatars that failed to load

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : render shared model row hints as native titles

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : fix badge guard for draft sidecars, keep parameter precision

hasBadges now counts draft sidecar badges, so a sidecar-only model still
renders. Billions keep one decimal for hub counts and stay bare for whole
values. Avatar failures track the org instead of the instance, and the
download progress bar no longer pulses while determinate.

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:18 +03:00
Aleksander Grygier 8664eaea30 ui : model download pipeline (#27959)
* ui : model download pipeline

Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.

Assisted-by: pi:GLM-5.3-Flash

* ui : route sidecar tag lookup through the sidecars util, validate the paused list

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:17 +03:00
Aleksander Grygier 4cfb6d1c75 ui : model memory-fit estimation (#27957)
* ui : model memory-fit estimation

Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.

Assisted-by: pi:GLM-5.3-Flash

* ui : cover the memory-fit and tool-use heuristics in tests

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:17 +03:00
Aleksander Grygier 9b43336114 ui : Hugging Face Hub data layer (#27947)
* ui : Hugging Face Hub data layer

Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.

Assisted-by: pi:GLM-5.3-Flash

* ui : strip provider tilde prefix from hub avatar urls

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : trim redundant comments in the HF data layer service

Per review: drop JSDoc that restates the method name and inline comments
that restate the code; keep only comments carrying non-obvious context.

Assisted-by: pi:zai-org/GLM-5.3-Flash

* ui : harden the HF data layer error typing, cover the helpers in tests

Carries the HTTP status on retryable fetch errors instead of matching the
message text. Marks expand-dependent catalog fields optional and documents
the data/models index pairing. Adds table tests for the pure helpers.

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:16 +03:00
Aleksander Grygier f653250407 ui : model id grammar for sidecars, quants and capability parsing (#27946)
* ui : model id grammar for sidecars, quants and capability parsing

Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.

Assisted-by: pi:GLM-5.3-Flash

* ui : escape sidecar tokens in the regex alternation

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:16 +03:00
Aleksander GrygierandPascal fa2bde5543 ui : type-safe API types, fetch helpers and download-ready models store plumbing (#29582)
* ui : type-safe API types, fetch helpers and download-ready models store plumbing

Assisted-by: pi:GLM-5.3-Flash

* ui : document the model list index pairing, fix an em-dash

Assisted-by: pi:zai-org/GLM-5.3-Flash

* Update tools/ui/src/lib/components/app/chat/index.ts

Co-authored-by: Pascal <admin@serveurperso.com>

---------

Co-authored-by: Pascal <admin@serveurperso.com>
2026-09-30 13:42:15 +03:00
Pascal 25747b08e7 openvino: serve GET_ROWS on a weight view from the base Constant (#28381)
* openvino: serve GET_ROWS on a weight view from the base Constant

Resolve view_src when collecting weight Constants so a view over a
quantized weight no longer becomes a dynamic typed Parameter, and fold
the row offset of the view into the gather indices instead of slicing
the dequantization subgraph.

* openvino: lift the quantized GET_ROWS view rejection

The supports_op rejection of a quantized src0 view with a nonzero
offset keeps the vs0 GET_ROWS cases of #28253 away from OpenVINO.
The weight view now resolves to the base Constant with the row offset
folded into the gather indices, so the rejection goes away.
b11284
2026-09-30 11:17:02 +02:00
Pascal db00347a4b ci : fix Fusion / metal by adding glm5-next to MTL.csv (#29712)
#27773 adds the glm5-next arch without its rows in the Metal fusion
baseline, so test-fusion --check fails on it. The rows come from
test-fusion --record on an M5 Max, and --check passes 270/270.
2026-09-30 09:50:23 +02:00
R0CKSTAR 272aad8b98 musa : define __CUDA_ARCH__ for device passes (#29508)
The MUSA vendor header never defined __CUDA_ARCH__, so every architecture
test in the shared ggml-cuda sources evaluated to 0.  Kernel bodies gated on
the architecture therefore compiled to nothing, for example the q8_0 -> f16
dequantization kernel in convert.cu, whose NO_DEVICE_CODE fallback expands to
an empty body in host code.

Report the newest architecture like the HIP backend does and exclude the
NVIDIA-only features explicitly, as they are not usable on MUSA.  Define it
for device passes only: CUB uses defined(__CUDA_ARCH__) to detect device
compilation, which is also how nvcc behaves.

Drop the now-redundant defined(__CUDA_ARCH__) checks in the architecture
comparisons: __CUDA_ARCH__ is undefined in host passes for CUDA and MUSA, and
HIP defines it for every pass, so both forms select the same branch.
b11282
2026-09-30 09:14:37 +02:00
Sigbjørn Skjæret 72db1e02ff ci : add models backend check (#29651)
* add models backend check

* t4-medium for faster build
2026-09-30 09:06:24 +02:00
Captain-Tripps 2a53ace3be SYCL: reduce tensor allreduce sync with pinned host buffers (#29604) b11280 2026-09-30 02:25:13 -04:00
649dcb1036 add GLM-5.3-Flash (GLM5-Next) support (#27773)
* Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx

* Add initial MTP support

* Merge branch optimizations. Reduce allocated compute buffer size, speed up long context decode, fla, and slight MTP improvements.

* Review driven changes, remove env vars, protect tensors

* Strip MTP for initial PR

* Clean up after mtp strip

* Clean up after mtp strip

* Update speculative.cpp

* Update llama-context.h

* Clean up after mtp strip

* Fix tokenizer ignore merges

* Improve quantization protection selection

* Refactor mhc helpers, graph base

* Lint Fixes

* Apply suggestions from code review

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

* Skip glm5-next in model saver, fix CRLF

* Skip glm5-next in sweep

* Remove T4 fallback

* Review cleanup

* Review suggestions

* Defer separate MTP gguf handling to MTP PR, drop filter

* Repad n_head_kv

* kpool init apply

* Order by descending score

* Drop guard

* read kpool from hparams, clarify kpool cache flags, remove kpool_build_state(nullptr)

* Add glm5-next support to model saver and add arch test fixture

* Review cleanup

* Kpool pooled caching clarify

* Add multi stream support

* Finish Rebase

* Sparse FA fir DSA prefill

* Const

* Update llama-model.cpp to fix rebase error

* gguf-py : merge tensor map entries for HC tensors

* model : use build_gdn_l2_norm in GLM5_NEXT implementation

* chore : remove trailing whitespace

* model : use new OP precision setting API in GLM5_NEXT implementation

* mtmd : use ggml_swiglu_clamp in GLM5V and apply the image token limit

The two clamps around swiglu_split are what ggml_swiglu_clamp already does,
so the clamp bounds collapse back to one value. GLM5V also never called
set_limit_image_tokens(), so --image-max-tokens had no effect.

Assisted-by: Claude Opus 5
(cherry picked from commit 46d18e12d422be4cc04a70e4a9a9e0168bb3d5b7)

* llama : keep the GLM5-Next k-pool layout across ubatches

The layout was rebuilt from a full cell scan on every ubatch. Pools are fixed
by the positions relative to the sequence's first one, so the layout now lives
on the memory and a ubatch only appends to it.

A sequence edit no longer stales every pooled key either, only the ones at or
after the edited position, which makes a tail seq_rm free. The pooling subgraph
is built unconditionally so the graph shape no longer changes every kpool
tokens, and the pool axis is folded into rows before soft_max, which otherwise
exceeds the CUDA gridDim.y limit past n_kv 262144.

Assisted-by: Claude Opus 5
(cherry picked from commit 5d1c40b93e17fddbf73b785efe43e0d02ccb3977)

* model : write the GLM5-Next recurrent rollback checkpoints

The conv state and the delta net state were only written to the live row, so a
rollback restored whatever the checkpoint rows happened to hold. Take the same
route as kimi-k3: build_recurrent_attn for the state, and write all K_rs conv
groups. That also drops a state view that assumed contiguous rows.

Enroll the arch in test-recurrent-state-rollback, which catches this under its
garbage-filled cache pass.

Assisted-by: Claude Opus 5
(cherry picked from commit 5ace37e86d5d448e83ef5dde5632c748185b18cd)

* llama: fix PR #27773 test-save-load-state restore failure

Clear the attention and indexer cache data after a failed hybrid state restore so restored NaNs cannot affect a later sequence.

Assisted-by: Codex

* llama: fix PR #27773 gpu-rocm graph reallocation

Reserve the full GLM5-Next pool capacity and dirty pool count. The gpu-rocm Test step aborts when n_new grows while the graph node count stays fixed; CUDA, Vulkan, Metal, and WebGPU checks report the same error.

Assisted-by: Codex

* llama : fix GLM5-Next k-pool layout staleness after edits and shared teardown

Two defects in the cross-ubatch k-pool layout added by the k-pool commit:

1. Wrong results. An edited sequence only rebuilt its pool layout when its cell
   count changed, so if the first ubatch after an edit added back exactly as many
   cells as were removed, the stale position-to-cell list survived. With a unified
   cache and more than one sequence, where another sequence takes the freed cells,
   the reused layout points at the wrong cells (CPU: large logit drift, CUDA: NaN).
   Rebuild whenever the sequence is stale, not only on a size mismatch.

2. Slowdown. "shared" mode was assumed to end only with an edit that forces a
   rebuild, but sharing also ends when the other sequence is removed. The survivor
   kept shared = true, pinning cache_safe off and re-pooling every pool on every
   ubatch (server trigger: n>1 completions with -kvu, via the seq_cp in
   copy_state_to). In seq_rm, if the layout has shared cells, stale every sequence
   so one rebuild re-derives sharing and cache_safe returns to 1.

Assisted-by: Claude Opus 5

* llama : fix build_attn_mha stream stride for non-contiguous q

build_attn_mha split the batch into streams with a stream stride of
q->nb[3]/n_stream. That only equals one stream's span, (ne[2]/n_stream)*nb[2],
when q is contiguous. GLM5-Next is nope-only, so it does not concat a rope part
and passes the permuted q_absorbed straight in, where nb[3] != ne[2]*nb[2]; the
stride was then n_head times too large and every stream s >= 1 read another
head's queries. Split-KV (-np N without --kv-unified) multi-stream prefill was
wrong for every stream past the first. Unified KV and decode were unaffected
(n_stream == 1, and decode takes the gather path). Other MLA models concat rope
so q is contiguous and the computed value is unchanged for them.

Compute the stride from the token dimension, which is identical for a
contiguous q.

Assisted-by: Claude Opus 5

* llama : re-derive GLM5-Next k-pool sharing on state_read/state_drop

The shared-cell teardown added to seq_rm (stale every sequence when the layout
has shared cells, so a survivor does not keep shared = true and pin cache_safe
off) was missing from the other paths that can free shared cells: state_read
and state_drop staled only the one sequence. Apply the same re-derivation there
and correct the comment that claimed sharing ends only via an edit or seq_rm.

Assisted-by: Claude Opus 5

* quant : drop duplicate GLM5-Next hc_ filter

The hc_ name filter was listed twice in the GLM5_NEXT protection block.

Assisted-by: Claude Opus 5

* glm5-next: scope K-pool cache access to indexed operations

* glm5-next: keep K-pool access in hybrid index memory

* glm5-next: keep mHC graph builders model-local

* glm5-next: mark only touched pools per ubatch

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Piotr Wilkin <ilintar@gmail.com>
b11279
2026-09-30 14:20:32 +08:00
Alessandro de Oliveira Faria (A.K.A.CABELO) 931351ea50 vendor: update BoringSSL to 0.20260929.0 (#29669) b11278 2026-09-30 12:44:32 +08:00
Aaron Teo eae11d2217 ggml-zdnn: impl buffer reset, fix memory leaks (#29637)
cont: fix code style

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
b11277
2026-09-30 12:21:41 +08:00
cqderekandMax Krasnyansky 19e28a2770 Hexagon f16 activation ops (#29209)
* hexagon: add F16 support for activation ops (SILU/GELU/GELU_QUICK/GEGLU/SWIGLU)

Widens ggml_hexagon_supported_activations() to accept F16 (src0/dst/src1
must agree on type), and adds F16 per-thread worker functions in
act-ops.c mirroring the existing F32 workers, backed by new HVX f16
kernels (hvx_sigmoid_f16_aa, hvx_tanh_f16_aa, hvx_mul_mul_f16_aa,
hvx_min_scalar_f16 family).

SILU, GELU, GELU_QUICK, GEGLU, and SWIGLU are verified correct on-device
(QRD8850) via test-backend-ops CPU-diffed correctness tests. SWIGLU_OAI's
F16 path is code-complete and builds clean on host + all 4 DSP arch
variants (v73/v75/v79/v81), but has no F16 test-case coverage in
test-backend-ops and is therefore unverified on-device in this change.

* hex-ops: align macros

* hex-ops: minor formatting

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
b11276
2026-09-29 14:59:13 -07:00
Xiang Chen a6ea155d3d gguf : reject tensor size that wraps after padding (#26979)
GGML_PAD(nbytes, alignment) wraps to 0 when nbytes is within
(alignment - 1) of SIZE_MAX, which silently bypassed the size
overflow guard in gguf_init_from_reader. Reject the tensor before
padding when nbytes + (alignment - 1) would overflow.

Adds a test-gguf handcrafted case (F32, ne = [4, 2^30-1, 2^30+1, 1])
whose ggml_nbytes = 2^64 - 16 lands in the wrap window. Fails on
master, passes with the guard.
b11275
2026-09-29 22:46:20 +02:00
Adrien Gallouët d3954b9324 ggml : check row bounds in get_rows_back (#29575)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11274
2026-09-29 22:44:41 +02:00
bosh 48de2a1bcb model : support classifier_pooling for rerankers (#29627)
* model : support classifier_pooling for ModernBERT rerankers

Assisted-by: Claude Opus 5.5

* model : read classifier pooling type in load_hparams

Write classifier.pooling_type from _try_set_pooling_type whenever the
config has classifier_pooling, and read it in
llama_model_base::load_hparams. ModernBERT falls back to mean when it
is unspecified.

Assisted-by: Claude Opus 5.5

* conversion : only accept cls and mean for classifier_pooling

Assisted-by: Claude Opus 5.5

* model : rename classifier_pooling_type to pooling_type_cls

Assisted-by: Claude Opus 5.5
b11273
2026-09-29 22:33:05 +02:00
Trivikram Reddy 7fee178464 hexagon: optimize concat op (#29673)
* hex-concat: reduce pkts in gather/transpose hot loop

gather directly into dst buffer, use special instruction for gather sync

* hex-concat: use fastdiv

replace calls to sw divide with fastpath

* hex-concat: optimize DMA-HVX pipeline and add transpose helpers
b11272
2026-09-29 12:52:10 -07:00
Pascal 6a2743f028 CUDA: bitonic argsort handles rows wider than one block (#28957)
Without CUB (HIP, MUSA) argsort ran the bitonic kernel with one thread
per padded column, so any row above 1024 entries launched an invalid
block configuration. Each thread now owns several columns, every stage
of the network runs all owned columns before the barrier, and the block
is capped at 1024 threads. Shared memory becomes the only bound, which
supports_op checks against the device instead of a fixed 1024.

Rows up to 1024 run the same work as before. Bit-exact with the CUB
path on rows of 2048.
2026-09-29 20:09:10 +02:00
thelittlefiremanandCarl Philipp Klemm 748d4225b9 ggml-cuda: HIP: optimize packed byte subtraction (__vsubss4 -> __vsub4) (#29478)
* ggml-cuda: HIP: optimize non-saturating packed byte subtraction (`__vsubss4`)

* CI: ignore 1 spilled vgpr in fattn_vec

---------

Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz>
b11270
2026-09-29 19:59:19 +02:00
Aaron Teo cee37ffea0 ci: add zdnn backend build but not test (#29541)
* ci: add zdnn backend build but not test

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* ci: attempt to run a ubuntu 26.04 container

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* ci: set shell to bash

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* ci: clean up comments

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* ggml-zdnn: fix compiler errors

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

* vendor: attempt to ignore warnings from vendored files

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>

---------

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
b11269
2026-09-29 20:45:14 +03:00
lhez 6dbbac4429 opencl: fix get_tensor for q5_K adreno gemm_nonshuffle kernel (#29555) b11268 2026-09-29 20:44:42 +03:00
Ruben Ortlam 5c200e0c8d vulkan: Tune GDN kernel, fix Intel performance (#29476)
* vulkan: tune GDN shader

* tune for Intel
b11267
2026-09-29 20:44:20 +03:00
Matt Corallo 83dd71f869 vulkan : Load F32 A matrix 2 at a time when its 2-aligned (#29254)
It turns out Intel doesn't particularly like loading F32s one at a
time and we already have the _2aliagned load logic in mul_mat_vec,
so here we use it.

While we do already check all the requirements to load elements 4
at a time across [B]F16 and F32, it turns out [B]F16 loading 4 at a
time is sometimes slower on very specific shapes on Intel BMG.
Loading 4 at a time is a bit faster on F32, but its not material
and I assume might be slower on other platforms.

Note that we also need to validate `a_offset` is 2-aligned in
`mul_mat_vec.comp`, which was missing in the original 2-way-load
patch.

Some selected speedups from `test-backend-ops perf` on a B60.

  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=1,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1704 runs -   767.17 us/run - 117.44 MFLOP/run - 153.08 GFLOPS
  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=1,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    2556 runs -   529.81 us/run - 117.44 MFLOP/run - 221.66 GFLOPS

  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=2,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1704 runs -   727.13 us/run - 234.88 MFLOP/run - 323.03 GFLOPS
  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=2,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    2130 runs -   528.84 us/run - 234.88 MFLOP/run - 444.15 GFLOPS

  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=3,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1704 runs -   702.19 us/run - 352.32 MFLOP/run - 501.74 GFLOPS
  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=3,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1988 runs -   532.14 us/run - 352.32 MFLOP/run - 662.08 GFLOPS

  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=4,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1278 runs -   919.50 us/run - 469.76 MFLOP/run - 510.89 GFLOPS
  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=4,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1917 runs -   543.69 us/run - 469.76 MFLOP/run - 864.03 GFLOPS

  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=5,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1197 runs -   892.12 us/run - 587.20 MFLOP/run - 658.21 GFLOPS
  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=5,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1881 runs -   575.17 us/run - 587.20 MFLOP/run -   1.02 TFLOPS

  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=8,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1498 runs -   716.40 us/run - 939.52 MFLOP/run -   1.31 TFLOPS
  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=8,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                    1819 runs -   576.36 us/run - 939.52 MFLOP/run -   1.63 TFLOPS

  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=512,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                   134 runs -  7467.09 us/run -  60.13 GFLOP/run -   8.05 TFLOPS
  MUL_MAT(type_a=f32,type_b=f32,m=4096,n=512,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0):                   134 runs -  7478.12 us/run -  60.13 GFLOP/run -   8.04 TFLOPS
b11266
2026-09-29 20:39:36 +03:00
Ankit Khandelwal 94a0ae3e72 vulkan: MOE aware mat_mul_id tile selection (#29182)
mut_mul_id selected its matmul tile with total token count.
For MoE dispatch grid the true N per workgroup is per-expert rows.
At pp128 on Sarvam 30B that is 6, not 128, so the picker took the l-tile for ~6 live rows.
Most workers in each group had nothing to do.
This wasted time. The slow part was 55% of the whole job.
b11265
2026-09-29 20:38:59 +03:00
da89bb3ccc ggml : fix c++ odr by properly using GGML_COMMON_DECL_CPP (#29504)
* fix c++ odr by properly using GGML_COMMON_DECL_CPP

* using actual field rather than macro

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

---------

Co-authored-by: XZiar <xziar@xziar.xziar>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b11264
2026-09-29 20:31:31 +03:00