11330 Commits
Author SHA1 Message Date
Aman GuptaandGeorgi Gerganov c061df1983 Qwen4Exp: add MTP (#29761)
* Qwen4Exp: add MTP

* remove has_state member, check via ctx_bufs being non-empty

* consistent naming + less verbose comments

* cont : clean-up recurrent memory

* cont : clean-up comments

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b11330
2026-10-01 14:13:27 +03:00
Aman Gupta 66e0c17ee1 llama: fix qwen4exp (#29751)
* llama: fix qwen4exp

* qwen4exp: keep kq_mask input the same shape
2026-10-01 14:13:27 +03:00
Oliver Simons 7677678503 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792)
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.

We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)
2026-10-01 13:05:29 +02:00
Xuan-Son Nguyen 552f18f912 mtmd: cap max_image to ubatch for non_causal models (#29773) b11327 2026-10-01 11:55:11 +02:00
Georgi Gerganov 5503b04b05 meta: clear inactive AllReduce shards with FILL, not SCALE (#29793) b11326 2026-10-01 12:43:57 +03:00
a u s t i n def4d406ae jinja : skip copying loop scope unless a loop filter needs it (#29776) b11325 2026-10-01 10:10:57 +02:00
Pranesh GonegandlaandPranesh Gonegandla 32dd62ee6d llama-mmap : avoid a second full-size copy of each tensor with direct-io (#29749)
Assisted-by: Claude

Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
b11324
2026-10-01 09:44:31 +02:00
uvos f11d642a27 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (#29572) b11323 2026-10-01 08:51:38 +03:00
Max Krasnyansky 3aa0ce9bca hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (#29785) b11322 2026-10-01 08:35:52 +03:00
Pradeep Rao b0aca3c653 BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (#29640)
* BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS

* AOCL-Blas : Add an AOCL-BLAS Quick Start and drop the fixed version path

* AOCL-BLAS doc : Note on ZenDNN
b11321
2026-10-01 08:35:01 +03:00
e-mon b8f96c3e82 common : add LLM-jp-4.1 Harmony dialect handler (#29681)
LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space
after every special token and parallel tool calls are separated by
<|end|>. The GPT-OSS handler rejects this output, so add a dedicated
handler, selected by the chat_format=llm-jp-harmony-v1 declaration in
the chat template.

Assisted-by: Claude Fable 5.1
b11320
2026-10-01 08:33:55 +03:00
lhez 3ec4df42d9 opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (#29698) b11319 2026-10-01 08:33:23 +03:00
Toki NasinandSigbjørn Skjæret db33d3cb89 vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (#29734)
* vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3

The original tokenizer configs for PLaMo-2 and PLaMo-3 have
`add_bos_token: true` and `add_eos_token: false` , but
_set_vocab_plamo() did not write the BOS/EOS metadata. The
PLAMO2 tokenizer path also ignored add_bos/add_eos during
tokenization.

Write the settings from tokenizer_config.json and honor them in
the PLAMO2 tokenization path. GGUFs without these keys keep the
previous behavior.

* Update conversion/base.py

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
b11318
2026-10-01 08:31:42 +03:00
Kushal Garg 7dad6db858 llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229)
* bench : fix verbosity filter to show GGML_LOG_ERROR (#28107)

* bench: remove dead variables
b11317
2026-10-01 08:20:00 +03:00
Georgi Gerganov 2232bc8b5f metal : use bf16 math for mxfp4 mul-mat (#29770) b11316 2026-10-01 08:19:24 +03:00
Marlon Paz 79625e056e llama-bench : fix docs (#29464)
* OoD documenatation for llama-bench

Signed-off-by: mairp <oec.valle.art@gmail.com>

* Unset default: auto

Signed-off-by: mairp <oec.valle.art@gmail.com>

---------

Signed-off-by: mairp <oec.valle.art@gmail.com>
2026-10-01 08:18:46 +03:00
CaramelizedCUDA 66bcc27706 docs : refresh CPU ops support matrix (#29666)
The committed docs/ops/CPU.csv is out of sync with the current
test-backend-ops suite: 11 ops with CPU support (COL2IM_1D,
MUL_MAT_HADAMARD, SWIGLU_CLAMP, MUL_MAT_W4A4/W4A8, MUL_MAT_ID_W4A4/W4A8,
DSV4_HC_COMB/PRE/POST, LIGHTNING_INDEXER) are missing entirely, and
many other ops have fewer test cases than the suite generates now.
docs/ops.md (which CI requires to match the CSVs) therefore
understates CPU support.

Regenerated with:
  test-backend-ops support -b CPU --output csv > docs/ops/CPU.csv
  scripts/create_ops_docs.py

Note: ADD1 now reads unsupported on CPU because ggml_add1 is
GGML_DEPRECATED and the suite no longer generates test cases for it;
the CPU implementation itself is still present.

Assisted-by: Xing
2026-10-01 08:17:51 +03:00
Kevin Hopper 10f340d1a2 model : re-enable -sm tensor for qwen4exp (#28569)
#27941 disabled -sm tensor for qwen4exp because test-llama-archs asserted on the
Meta device once the fixture carried a PLE layer:
GGML_ASSERT(ggml_backend_buffer_is_meta(tensor->buffer)) at ggml-backend-meta.cpp:476.

With host-resident embeddings the PLE gather is a CPU node and hc_init (the REPEAT
that fans the embedding out to the hc streams) was first reached through layer 0's
PLE path, after that gather. ggml_backend_sched_split_graph pass 2 expands a device
assignment upwards only until it meets a CPU node, so the REPEAT stayed on the CPU
and the later reshape of hc_init inside the meta split viewed a host-resident node.

Expanding hc_init right after it is built puts the REPEAT directly before the first
device node, where pass 2 assigns it; the embedding reshape stays in the CPU split
and is copied in as a split input, as in deepseek4.
b11313
2026-10-01 08:16:49 +03:00
Masashi Yoshimura 0c1e57098b webgpu: fix SSM_SCAN binding aliasing (#29750) b11312 2026-10-01 11:11:48 +09:00
Adrien Gallouët f7b384c1e5 ggml-opencl : replace alloca() with std::vector (#29765)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11311
2026-09-30 23:59:46 +02:00
R0CKSTAR f872b59112 cuda: guard the iq4_nl dequantize row kernel against short rows (#29683)
dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than that (an IQ4_NL row is only guaranteed to be a multiple of QK4_NL). Threads whose 32-value sub-block starts at or past k currently read and write past the end of the row. Skip those sub-blocks; for rows that are a multiple of QK_K the check never fires.
b11310
2026-09-30 22:33:22 +02:00
Ehsan BateniandMax Krasnyansky a4d880fd5c Hexagon: optimize ALLREDUCE with support for safe scatter mode (#29757)
* hex-allreduce: add support for safe scatter mode

* hex-allreduce: pare down excessive comments

* hex-allreduce: re-write to remove register spills

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
b11309
2026-09-30 12:56:44 -07:00
pr3pony feb9a3d6de args: fix cli download mmproj arg (#28977)
* tests: add tests for cli download arg parsing

* args: fix cli download mmproj arg
b11308
2026-09-30 20:45:39 +02:00
Hrishith Thadicherla 4453b535fd llama : preserve original batch order for speculative decoding layer inputs (#29019)
* llama: preserve original batch order for layer inputs

Assisted-by: Codex

* tests: cover layer-input order across KV layouts

Assisted-by: Codex

* tests: exercise layer-input ordering on CUDA devices

Assisted-by: Codex

* llama: make layer input reordering compatible with tensor split

Copy each microbatch tensor from offset zero and restore original row order after synchronization. Extend the layer-input regression to cover tensor split and repeated reads and decodes.

Assisted-by: Codex

* llama: restore token order for unmasked NextN embeddings

Use the original-token mapping for unmasked NextN rows, including when
layer-input capture is disabled. Keep masked NextN rows on the logits
output mapping and preserve offset-zero tensor copies.

Extend the existing regression to cover NextN alone, combined layer
capture, and masked outputs with repeated decodes and getters.

Validation: all 256 CPU/CUDA/tensor configurations pass. Qwen3.8-27B
Q4_K_M MTP completes MT-Bench at concurrency 16 before and after.

Assisted-by: Codex

* ggml: fix WebGPU reservation and OpenVINO hidden-state capture

Reserve WebGPU vector attention scratch across batch sizes and refresh reservations when NextN capture settings change. Preserve requested OpenVINO outputs, dynamic shapes, sequence counts, and current graph bindings.

Extend existing WebGPU regression coverage and enable strict allocation checks.

Assisted-by: Codex

* llama: defer regression test and backend fixes to follow-ups

Keep this PR focused on restoring token order for layer inputs and unmasked NextN embeddings. Remove the added regression test, OpenVINO and WebGPU changes, and the separate NextN reservation change.

Assisted-by: Codex

* llama: keep n_embd declaration in its original position

Assisted-by: Codex

* llama : pass token count to layer input extraction

Assisted-by: Codex

* llama : name original batch indices batch_idxs

Assisted-by: Codex

* llama : name extracted embedding indices embd_batch_idxs

Assisted-by: Codex

* llama : tag target embedding reordering

Assisted-by: Codex

* llama : tag extraction and name the index capture flag

Assisted-by: Codex
b11307
2026-09-30 21:17:40 +03:00
Sihan Yu 4f31296a90 test-llama-archs : toggle causal_attn to catch graph shape changes (#29724)
After the device decode, flip causal_attn off, decode n_ubatch/2 then
n_ubatch tokens. Both have the same node count, so a shape that depends
on the flag makes the second reallocate at an unchanged graph size,
which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
b11306
2026-09-30 20:39:53 +03:00
b016f461be convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324)
* convert: fix LoRA conversion crash for Qwen3.5 V-head reorder

_reorder_v_heads does reshape+permute+reshape to reorder V heads from
grouped to tiled order.  LoraTorchTensor.reshape() cannot split its
row dimension (A matrix), so converting Qwen3.5 LoRA adapters that
target out_proj crashes with NotImplementedError.

Fix: detect LoRA tensors and apply the equivalent index permutation
directly — column reorder (dim=last) permutes A's columns, row
reorder (dim=0) permutes B's rows.  This is mathematically identical:
  (B @ A)[:, perm] == B @ A[:, perm]
  (B @ A)[perm, :] == B[perm, :] @ A

Verified: both paths produce exactly zero diff against the full-tensor
reorder on random (rank=32, 4096×4096) matrices.

Fixes #21125

Signed-off-by: Radu Swigler <radu@swigler.com>

* convert: add ty: ignore for hasattr-guarded LoRA call

Assisted-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix comment

* nowrap

---------

Signed-off-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Assisted-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-30 19:30:21 +02:00
Xuan-Son Nguyen 81ff93ea1d llama: properly handle KV on training (#28520)
* llama: properly handle KV on training

* improve
b11304
2026-09-30 18:09:08 +02:00
Xuan-Son Nguyen 60e9cf7a7b batch: migrate the rest of examples to llama_batch_ext (#29601)
* migrate the rest

* test-thread-safety

* rm common_batch_staged
b11303
2026-09-30 18:08:43 +02:00
Pascal 05af0d2b13 glm5-next: give dead indexer slots unique scatter rows (#29745)
The sparse indexer mask is built with a set_rows scatter. Padded pools,
absent sequences and missing tail cells all pointed to the same n_kv
sentinel row, and invisible pools picked by top_k to fill the selection
overlap the tail cells of the token, so several CPU threads wrote the
same element (ThreadSanitizer data race in the sanitize CI).

Allocate the slot mask for both selection paths and route every dead
slot to its own dump row n_kv + slot. Live slots address disjoint cells,
so the scatter indices of a token are unique.
b11302
2026-09-30 17:10:48 +02:00
Daniel Kuts 2149c00f44 ggml/gguf : fix integer overflow (#29384)
* ggml: fix integer overflow guard for zero-element tensors

* ggml: validate number of elements in tensor to prevent integer overflow

* ggml: fix error print
b11301
2026-09-30 17:59:00 +03:00
Vishal SinghandVishal Singh 876c75b1f6 codeowners : remove former ZenDNN owner (#29747)
Co-authored-by: Vishal Singh <numeric-id+vishalMCE@users.noreply.github.com>
2026-09-30 22:38:21 +08:00
Pascal b04642061d cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (#29722)
* cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast

On Windows the simple input reader sends CTRL_C_EVENT to every process
attached to the console when stdin reaches EOF, killing unrelated
processes such as a supervising agent. The CLI only stopped on EOF
because of that self inflicted SIGINT; on POSIX, and with the advanced
reader, it spins forever printing prompts.

Drop the broadcast so both platforms just return an empty read, and
treat an empty read as EOF in the chat loop and the model selection,
since a submitted line always ends with a newline.

* cli: keep the newline of a trailing "/" and stop mtmd-cli on EOF

A lone "/" came back as an empty read and was taken for EOF, and
mtmd-cli only stopped on EOF through the removed broadcast.
b11299
2026-09-30 16:24:46 +02:00
Georgi GerganovandSigbjørn Skjæret 22bdcc4cdd mimo : support dflash (convert + feature extraction) (#29650)
* convert : update to support dflash

* cont : fix

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
b11298
2026-09-30 17:04:35 +03:00
Sigbjørn Skjæret ca2e2037b6 jinja : support coerced array attributes (#29574)
* support coerced array attributes

* add tests
b11297
2026-09-30 15:33:26 +02:00
Adrien Gallouët bdeb855b30 ggml-et : remove useless alloca() (#29663)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-30 15:14:45 +02:00
Pascal 3b3d022b82 ci : fix Models Backend Check by shortening the hrm_text fixture (#29744)
The fixture recycles its two blocks over 8 cache slots, so the fp16
error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU
jobs of Models Backend. Two l-cycles keep every branch of the cycle
loop and halve the error.
b11295
2026-09-30 15:03:38 +02:00
185103dcf5 llama: llama_prefetch_rows (#29599)
* llama: llama_prefetch_rows

* llama: support row prefetch on Windows

Apply the Windows port contributed by @praneshgo unchanged.

Source: https://github.com/ggml-org/llama.cpp/pull/29599#issuecomment-5887721014

* avoid exposing llama-mmap in model code, route via llama-impl

* add windows check, only prefetch in lazy mode

* cont : clean-up

* cont : fix build

* cont : clarify padding token for gemma4

---------

Co-authored-by: Pranesh Gonegandla <pranesh.iitp@gmail.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b11294
2026-09-30 20:27:18 +08:00
Aman Gupta 2090f60f0b ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA) (#29675)
* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)

* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops

Assisted-by: Claude Opus 5.5

* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build

* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
b11293
2026-09-30 20:24:59 +08:00
Pascal 90c908d06d cpu: accept BF16 in src1 of mul_mat (#28937)
* cpu: accept BF16 in src1 of mul_mat

ggml_conv_1d_dw builds its im2col in F32 when the kernel is BF16, then
calls ggml_mul_mat(im2col, kernel), which puts F32 in src0 and BF16 in
src1. The CPU backend refused that combination, so it was reported as
unsupported on every backend and never compared against anything.

Widen BF16 into the F32 work buffer, next to the existing packing of F32
into vec_dot_type. This is the arithmetic the Metal mat vec kernel
already uses, both operands promoted to float and accumulated in float,
so the two agree exactly rather than approximately.

Cover it with a conv_1d_dw test over F32, F16 and BF16 kernels, plus
three mul_mat cases with BF16 in src1.

* vulkan: reject BF16 in src1 of mul_mat unless src0 is BF16

supports_op only checked the src1 type for non contiguous tensors, so
a contiguous BF16 src1 was accepted and the pipeline lookup asserted.
The only BF16 src1 path is the BF16 x BF16 multiply, every other src0
type now reports the op as unsupported and the scheduler keeps it on
the CPU.

The BF16 kernel case of the conv_1d_dw test needs the f32 x bf16
mat vec variants of the Metal backend, which land separately.
b11292
2026-09-30 14:14:45 +02:00
Daniel Bevenius 8df332de1b model-conversion : add --add-bos to run org model script (#29558)
This commit adds an optional --add-bos token command line option to the
run-org-model.py script.

The motivation for this is that there are models, for example Gemma4,
that explicitely set the add_bos value to true in llama-vocab.cpp even
if the original model does not set this value to True.

It would be nice to be able to force the models to agree on the bos
token so that logit verification can proceed.

Refs: https://github.com/ggml-org/llama.cpp/pull/21500
2026-09-30 12:52:02 +02:00
Aleksander Grygier 4a096b8ff6 ui : shared model display primitives (#29644)
* ui : shared model display primitives

Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.

Assisted-by: pi:GLM-5.3-Flash

* ui : remember hub avatars that failed to load

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : render shared model row hints as native titles

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : fix badge guard for draft sidecars, keep parameter precision

hasBadges now counts draft sidecar badges, so a sidecar-only model still
renders. Billions keep one decimal for hub counts and stay bare for whole
values. Avatar failures track the org instead of the instance, and the
download progress bar no longer pulses while determinate.

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:18 +03:00
Aleksander Grygier 8664eaea30 ui : model download pipeline (#27959)
* ui : model download pipeline

Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.

Assisted-by: pi:GLM-5.3-Flash

* ui : route sidecar tag lookup through the sidecars util, validate the paused list

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:17 +03:00
Aleksander Grygier 4cfb6d1c75 ui : model memory-fit estimation (#27957)
* ui : model memory-fit estimation

Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.

Assisted-by: pi:GLM-5.3-Flash

* ui : cover the memory-fit and tool-use heuristics in tests

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:17 +03:00
Aleksander Grygier 9b43336114 ui : Hugging Face Hub data layer (#27947)
* ui : Hugging Face Hub data layer

Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.

Assisted-by: pi:GLM-5.3-Flash

* ui : strip provider tilde prefix from hub avatar urls

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : trim redundant comments in the HF data layer service

Per review: drop JSDoc that restates the method name and inline comments
that restate the code; keep only comments carrying non-obvious context.

Assisted-by: pi:zai-org/GLM-5.3-Flash

* ui : harden the HF data layer error typing, cover the helpers in tests

Carries the HTTP status on retryable fetch errors instead of matching the
message text. Marks expand-dependent catalog fields optional and documents
the data/models index pairing. Adds table tests for the pure helpers.

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:16 +03:00
Aleksander Grygier f653250407 ui : model id grammar for sidecars, quants and capability parsing (#27946)
* ui : model id grammar for sidecars, quants and capability parsing

Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.

Assisted-by: pi:GLM-5.3-Flash

* ui : escape sidecar tokens in the regex alternation

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:16 +03:00
Aleksander GrygierandPascal fa2bde5543 ui : type-safe API types, fetch helpers and download-ready models store plumbing (#29582)
* ui : type-safe API types, fetch helpers and download-ready models store plumbing

Assisted-by: pi:GLM-5.3-Flash

* ui : document the model list index pairing, fix an em-dash

Assisted-by: pi:zai-org/GLM-5.3-Flash

* Update tools/ui/src/lib/components/app/chat/index.ts

Co-authored-by: Pascal <admin@serveurperso.com>

---------

Co-authored-by: Pascal <admin@serveurperso.com>
2026-09-30 13:42:15 +03:00
Pascal 25747b08e7 openvino: serve GET_ROWS on a weight view from the base Constant (#28381)
* openvino: serve GET_ROWS on a weight view from the base Constant

Resolve view_src when collecting weight Constants so a view over a
quantized weight no longer becomes a dynamic typed Parameter, and fold
the row offset of the view into the gather indices instead of slicing
the dequantization subgraph.

* openvino: lift the quantized GET_ROWS view rejection

The supports_op rejection of a quantized src0 view with a nonzero
offset keeps the vs0 GET_ROWS cases of #28253 away from OpenVINO.
The weight view now resolves to the base Constant with the row offset
folded into the gather indices, so the rejection goes away.
b11284
2026-09-30 11:17:02 +02:00
Pascal db00347a4b ci : fix Fusion / metal by adding glm5-next to MTL.csv (#29712)
#27773 adds the glm5-next arch without its rows in the Metal fusion
baseline, so test-fusion --check fails on it. The rows come from
test-fusion --record on an M5 Max, and --check passes 270/270.
2026-09-30 09:50:23 +02:00
R0CKSTAR 272aad8b98 musa : define __CUDA_ARCH__ for device passes (#29508)
The MUSA vendor header never defined __CUDA_ARCH__, so every architecture
test in the shared ggml-cuda sources evaluated to 0.  Kernel bodies gated on
the architecture therefore compiled to nothing, for example the q8_0 -> f16
dequantization kernel in convert.cu, whose NO_DEVICE_CODE fallback expands to
an empty body in host code.

Report the newest architecture like the HIP backend does and exclude the
NVIDIA-only features explicitly, as they are not usable on MUSA.  Define it
for device passes only: CUB uses defined(__CUDA_ARCH__) to detect device
compilation, which is also how nvcc behaves.

Drop the now-redundant defined(__CUDA_ARCH__) checks in the architecture
comparisons: __CUDA_ARCH__ is undefined in host passes for CUDA and MUSA, and
HIP defines it for every pass, so both forms select the same branch.
b11282
2026-09-30 09:14:37 +02:00
Sigbjørn Skjæret 72db1e02ff ci : add models backend check (#29651)
* add models backend check

* t4-medium for faster build
2026-09-30 09:06:24 +02:00