Commit Graph
11350 Commits
Author SHA1 Message Date
Oliver Simons aa742bb620 More FP8 conversion fixes 2026-10-01 17:59:18 +02:00
Oliver Simons 471b26dd4c Emit warnings when not fusing FP8 during conversion 2026-10-01 17:59:18 +02:00
Oliver Simons e703584f32 Avoid fusing FP8 for --fuse-gate-up-exps as we may have different ws 2026-10-01 17:59:18 +02:00
Oliver Simons b2860f80b1 Fix conversion of fused qkvz & FP8 for Qwen3-Next 2026-10-01 17:59:18 +02:00
Oliver Simons 5f5aedd531 Disallow QKV-fusion for FP8 temporarily 2026-10-01 17:59:18 +02:00
Oliver Simons 7ba386e041 Fix and test gguf-py behavior for FP8_E4M3 2026-10-01 17:59:18 +02:00
Oliver Simons 1c4cde0e4b Allow tensor-strategy for FP8-only compressed-tensor ckpts 2026-10-01 17:59:18 +02:00
Oliver Simons 19361354a1 Allow remaining FP8 tensors to be dequantized during conversion 2026-10-01 17:59:18 +02:00
Oliver Simons 9e12cdadfe Fix conversion of https://huggingface.co/RedHatAI/Qwen3.8-27B-NVFP4
1. Per row FP8 wrongly fell to NVFP4 conversion, as NVFP4 checked on
   dims alone. Also check on dtype for FP8
2. Need to reshape FP8 QKV projections scales in the same way that
   weights are reshaped
2026-10-01 17:59:18 +02:00
Oliver Simons 1d7db01780 Materialize input_scales also for FP8 when converting 2026-10-01 17:59:18 +02:00
Oliver Simons eff0a3dbf8 Fix python type-checker 2026-10-01 17:59:18 +02:00
Oliver Simons 3932b244e9 WIP OCP FP8 E4M3 support 2026-10-01 17:59:18 +02:00
Yiwei ShaoandMax Krasnyansky dcd387a412 hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (#29685)
* hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA

* hex-cpy: various fixes on top of the concat optimizations

Removed CONCAT_DMA_MIN_ROW logic, it was broken with 64-bit DMA.
While it's kinda silly to use DMA for tiny stuff if that tensor gets mapped to an extended buffer the only way to read it is DMA.

Added missing dma_queue_flush() calls.

Added additional guards for conditions we don't support.

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
b11338
2026-10-01 08:38:37 -07:00
Sam Malayek d775ebf363 server: return HTTP 400 for invalid embedding requests (#29060) b11337 2026-10-01 16:52:43 +02:00
Yu Chengye 2b36825cbc convert : write Gemma embedding scale for DFlash drafts (#29802)
* convert : write Gemma embedding scale for DFlash drafts

A DFlash draft shares the target's token embeddings. Gemma scales them by sqrt(hidden_size) in the forward pass, and the draft config does not state that scale, so the converted draft read unscaled embeddings.

Take the scale from the target config when the draft config has none.

Assisted-by: Claude

* convert : check with get_model_architecture for gemma models
2026-10-01 16:34:32 +02:00
42d958167a cuda : route sm70 to the Turing MMVQ nwarps table (#29753)
* cuda : route sm70 to the Turing MMVQ nwarps table

Volta (sm_70) has no MMVQ parameter table of its own and falls through
to GENERIC, which launches K-quant batch-1 decode (ncols_dst == 1) at
nwarps=4. sm_70 shares TURING's tuning: the K-quant vec_dot prefers
nwarps=2 there. Route sm_70 to the existing MMVQ_PARAMETERS_TURING
table in both the device and the host table selector.

Measured on one Tesla V100 32GB PCIe (PG500-216, driver 580.178.04,
CUDA 12.0.140) with Qwen3.8-27B Q4_K_M, tg128, interleaved A/B in 6
ABBA blocks with paired per-block deltas: +1.091 t/s = +3.17 %
(t = +49.0, all six per-block deltas positive); perplexity
bit-identical (6.3697 +/- 0.04066 both builds, wiki.test.raw). The
patched build's K-quant mul_mat_vec_q kernels launch at nwarps=2
(cubin EIATTR_MAX_THREADS) while Q4_0/Q8_0 stay at nwarps=4, and the
same measurement on the September master base gave +3.84 % (t = 85).

The tuning originates from the V100-focused fork anyei/llamacpp-v100
(MIT), commit b912d1b1e, which carries a dedicated
MMVQ_PARAMETERS_VOLTA table; a cubin-level comparison confirmed that
routing sm_70 to the existing TURING table is equivalent for the
K-quant batch-1 path this change affects, so this is the minimal
2-line form. https://github.com/anyei/llamacpp-v100/commit/b912d1b1e

Original-patch-by: anyei <angelyoelroblesmercedes@gmail.com>

* Update ggml/src/ggml-cuda/mmvq.cu

---------

Co-authored-by: tkittich <tkittich@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
b11335
2026-10-01 15:21:26 +02:00
Mike van LammerenandNiklas Wenzel 13b4d7135a metal : release temporary private transfer buffers (#29777)
* metal : release temporary private transfer buffers

Assisted-by: OpenAI Codex

* metal : fix order and formatting

---------

Co-authored-by: Niklas Wenzel <dev@nikwen.de>
b11334
2026-10-01 16:17:23 +03:00
Masashi Yoshimura 4b1622afb7 webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (#29358) b11333 2026-10-01 22:09:07 +09:00
Georgi Gerganov 869034b4bb llama : fix invalid assert in recurrent memory (#29799) b11332 2026-10-01 14:49:03 +03:00
ynankaniandJohannes Gäßler b56f34ab13 CUDA: Handle compute type for NVFP4 on cublass path (#29173)
* CUDA: Handle compute type for NVFP4 on cublass path

Signed-off-by: ynankani <ynankani@nvidia.com>

* Use BF16 compute type for quantized models if HW allows

Signed-off-by: ynankani <ynankani@nvidia.com>

* Set acc prec to bf16 for nvfp4 as it needs atleast bf16 range

Signed-off-by: ynankani <ynankani@nvidia.com>

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* preserve op_params for per-expert matmul

Signed-off-by: ynankani <ynankani@nvidia.com>

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
b11331
2026-10-01 16:53:52 +05:30
Aman GuptaandGeorgi Gerganov c061df1983 Qwen4Exp: add MTP (#29761)
* Qwen4Exp: add MTP

* remove has_state member, check via ctx_bufs being non-empty

* consistent naming + less verbose comments

* cont : clean-up recurrent memory

* cont : clean-up comments

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b11330
2026-10-01 14:13:27 +03:00
Aman Gupta 66e0c17ee1 llama: fix qwen4exp (#29751)
* llama: fix qwen4exp

* qwen4exp: keep kq_mask input the same shape
2026-10-01 14:13:27 +03:00
Oliver Simons 7677678503 CUDA: Make CCCL configurable + pin it to 3.4.3 for CI jobs (#29792)
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.

We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)
2026-10-01 13:05:29 +02:00
Xuan-Son Nguyen 552f18f912 mtmd: cap max_image to ubatch for non_causal models (#29773) b11327 2026-10-01 11:55:11 +02:00
Georgi Gerganov 5503b04b05 meta: clear inactive AllReduce shards with FILL, not SCALE (#29793) b11326 2026-10-01 12:43:57 +03:00
a u s t i n def4d406ae jinja : skip copying loop scope unless a loop filter needs it (#29776) b11325 2026-10-01 10:10:57 +02:00
Pranesh GonegandlaandPranesh Gonegandla 32dd62ee6d llama-mmap : avoid a second full-size copy of each tensor with direct-io (#29749)
Assisted-by: Claude

Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
b11324
2026-10-01 09:44:31 +02:00
uvos f11d642a27 HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (#29572) b11323 2026-10-01 08:51:38 +03:00
Max Krasnyansky 3aa0ce9bca hex-workqueue: fix race condition in seqn getting out of sync with idx_read/write (#29785) b11322 2026-10-01 08:35:52 +03:00
Pradeep Rao b0aca3c653 BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS (#29640)
* BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS

* AOCL-Blas : Add an AOCL-BLAS Quick Start and drop the fixed version path

* AOCL-BLAS doc : Note on ZenDNN
b11321
2026-10-01 08:35:01 +03:00
e-mon b8f96c3e82 common : add LLM-jp-4.1 Harmony dialect handler (#29681)
LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space
after every special token and parallel tool calls are separated by
<|end|>. The GPT-OSS handler rejects this output, so add a dedicated
handler, selected by the chat_format=llm-jp-harmony-v1 declaration in
the chat template.

Assisted-by: Claude Fable 5.1
b11320
2026-10-01 08:33:55 +03:00
lhez 3ec4df42d9 opencl: mark vec subgroup bcast as supproted for Adreno E17 compiler (#29698) b11319 2026-10-01 08:33:23 +03:00
Toki NasinandSigbjørn Skjæret db33d3cb89 vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3 (#29734)
* vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3

The original tokenizer configs for PLaMo-2 and PLaMo-3 have
`add_bos_token: true` and `add_eos_token: false` , but
_set_vocab_plamo() did not write the BOS/EOS metadata. The
PLAMO2 tokenizer path also ignored add_bos/add_eos during
tokenization.

Write the settings from tokenizer_config.json and honor them in
the PLAMO2 tokenization path. GGUFs without these keys keep the
previous behavior.

* Update conversion/base.py

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
b11318
2026-10-01 08:31:42 +03:00
Kushal Garg 7dad6db858 llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229)
* bench : fix verbosity filter to show GGML_LOG_ERROR (#28107)

* bench: remove dead variables
b11317
2026-10-01 08:20:00 +03:00
Georgi Gerganov 2232bc8b5f metal : use bf16 math for mxfp4 mul-mat (#29770) b11316 2026-10-01 08:19:24 +03:00
Marlon Paz 79625e056e llama-bench : fix docs (#29464)
* OoD documenatation for llama-bench

Signed-off-by: mairp <oec.valle.art@gmail.com>

* Unset default: auto

Signed-off-by: mairp <oec.valle.art@gmail.com>

---------

Signed-off-by: mairp <oec.valle.art@gmail.com>
2026-10-01 08:18:46 +03:00
CaramelizedCUDA 66bcc27706 docs : refresh CPU ops support matrix (#29666)
The committed docs/ops/CPU.csv is out of sync with the current
test-backend-ops suite: 11 ops with CPU support (COL2IM_1D,
MUL_MAT_HADAMARD, SWIGLU_CLAMP, MUL_MAT_W4A4/W4A8, MUL_MAT_ID_W4A4/W4A8,
DSV4_HC_COMB/PRE/POST, LIGHTNING_INDEXER) are missing entirely, and
many other ops have fewer test cases than the suite generates now.
docs/ops.md (which CI requires to match the CSVs) therefore
understates CPU support.

Regenerated with:
  test-backend-ops support -b CPU --output csv > docs/ops/CPU.csv
  scripts/create_ops_docs.py

Note: ADD1 now reads unsupported on CPU because ggml_add1 is
GGML_DEPRECATED and the suite no longer generates test cases for it;
the CPU implementation itself is still present.

Assisted-by: Xing
2026-10-01 08:17:51 +03:00
Kevin Hopper 10f340d1a2 model : re-enable -sm tensor for qwen4exp (#28569)
#27941 disabled -sm tensor for qwen4exp because test-llama-archs asserted on the
Meta device once the fixture carried a PLE layer:
GGML_ASSERT(ggml_backend_buffer_is_meta(tensor->buffer)) at ggml-backend-meta.cpp:476.

With host-resident embeddings the PLE gather is a CPU node and hc_init (the REPEAT
that fans the embedding out to the hc streams) was first reached through layer 0's
PLE path, after that gather. ggml_backend_sched_split_graph pass 2 expands a device
assignment upwards only until it meets a CPU node, so the REPEAT stayed on the CPU
and the later reshape of hc_init inside the meta split viewed a host-resident node.

Expanding hc_init right after it is built puts the REPEAT directly before the first
device node, where pass 2 assigns it; the embedding reshape stays in the CPU split
and is copied in as a split input, as in deepseek4.
b11313
2026-10-01 08:16:49 +03:00
Masashi Yoshimura 0c1e57098b webgpu: fix SSM_SCAN binding aliasing (#29750) b11312 2026-10-01 11:11:48 +09:00
Adrien Gallouët f7b384c1e5 ggml-opencl : replace alloca() with std::vector (#29765)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11311
2026-09-30 23:59:46 +02:00
R0CKSTAR f872b59112 cuda: guard the iq4_nl dequantize row kernel against short rows (#29683)
dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than that (an IQ4_NL row is only guaranteed to be a multiple of QK4_NL). Threads whose 32-value sub-block starts at or past k currently read and write past the end of the row. Skip those sub-blocks; for rows that are a multiple of QK_K the check never fires.
b11310
2026-09-30 22:33:22 +02:00
Ehsan BateniandMax Krasnyansky a4d880fd5c Hexagon: optimize ALLREDUCE with support for safe scatter mode (#29757)
* hex-allreduce: add support for safe scatter mode

* hex-allreduce: pare down excessive comments

* hex-allreduce: re-write to remove register spills

---------

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
b11309
2026-09-30 12:56:44 -07:00
pr3pony feb9a3d6de args: fix cli download mmproj arg (#28977)
* tests: add tests for cli download arg parsing

* args: fix cli download mmproj arg
b11308
2026-09-30 20:45:39 +02:00
Hrishith Thadicherla 4453b535fd llama : preserve original batch order for speculative decoding layer inputs (#29019)
* llama: preserve original batch order for layer inputs

Assisted-by: Codex

* tests: cover layer-input order across KV layouts

Assisted-by: Codex

* tests: exercise layer-input ordering on CUDA devices

Assisted-by: Codex

* llama: make layer input reordering compatible with tensor split

Copy each microbatch tensor from offset zero and restore original row order after synchronization. Extend the layer-input regression to cover tensor split and repeated reads and decodes.

Assisted-by: Codex

* llama: restore token order for unmasked NextN embeddings

Use the original-token mapping for unmasked NextN rows, including when
layer-input capture is disabled. Keep masked NextN rows on the logits
output mapping and preserve offset-zero tensor copies.

Extend the existing regression to cover NextN alone, combined layer
capture, and masked outputs with repeated decodes and getters.

Validation: all 256 CPU/CUDA/tensor configurations pass. Qwen3.8-27B
Q4_K_M MTP completes MT-Bench at concurrency 16 before and after.

Assisted-by: Codex

* ggml: fix WebGPU reservation and OpenVINO hidden-state capture

Reserve WebGPU vector attention scratch across batch sizes and refresh reservations when NextN capture settings change. Preserve requested OpenVINO outputs, dynamic shapes, sequence counts, and current graph bindings.

Extend existing WebGPU regression coverage and enable strict allocation checks.

Assisted-by: Codex

* llama: defer regression test and backend fixes to follow-ups

Keep this PR focused on restoring token order for layer inputs and unmasked NextN embeddings. Remove the added regression test, OpenVINO and WebGPU changes, and the separate NextN reservation change.

Assisted-by: Codex

* llama: keep n_embd declaration in its original position

Assisted-by: Codex

* llama : pass token count to layer input extraction

Assisted-by: Codex

* llama : name original batch indices batch_idxs

Assisted-by: Codex

* llama : name extracted embedding indices embd_batch_idxs

Assisted-by: Codex

* llama : tag target embedding reordering

Assisted-by: Codex

* llama : tag extraction and name the index capture flag

Assisted-by: Codex
b11307
2026-09-30 21:17:40 +03:00
Sihan Yu 4f31296a90 test-llama-archs : toggle causal_attn to catch graph shape changes (#29724)
After the device decode, flip causal_attn off, decode n_ubatch/2 then
n_ubatch tokens. Both have the same node count, so a shape that depends
on the flag makes the second reallocate at an unchanged graph size,
which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
b11306
2026-09-30 20:39:53 +03:00
b016f461be convert : fix LoRA conversion crash for Qwen3.5 V-head reorder (#28324)
* convert: fix LoRA conversion crash for Qwen3.5 V-head reorder

_reorder_v_heads does reshape+permute+reshape to reorder V heads from
grouped to tiled order.  LoraTorchTensor.reshape() cannot split its
row dimension (A matrix), so converting Qwen3.5 LoRA adapters that
target out_proj crashes with NotImplementedError.

Fix: detect LoRA tensors and apply the equivalent index permutation
directly — column reorder (dim=last) permutes A's columns, row
reorder (dim=0) permutes B's rows.  This is mathematically identical:
  (B @ A)[:, perm] == B @ A[:, perm]
  (B @ A)[perm, :] == B[perm, :] @ A

Verified: both paths produce exactly zero diff against the full-tensor
reorder on random (rank=32, 4096×4096) matrices.

Fixes #21125

Signed-off-by: Radu Swigler <radu@swigler.com>

* convert: add ty: ignore for hasattr-guarded LoRA call

Assisted-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix comment

* nowrap

---------

Signed-off-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Assisted-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-30 19:30:21 +02:00
Xuan-Son Nguyen 81ff93ea1d llama: properly handle KV on training (#28520)
* llama: properly handle KV on training

* improve
b11304
2026-09-30 18:09:08 +02:00
Xuan-Son Nguyen 60e9cf7a7b batch: migrate the rest of examples to llama_batch_ext (#29601)
* migrate the rest

* test-thread-safety

* rm common_batch_staged
b11303
2026-09-30 18:08:43 +02:00
Pascal 05af0d2b13 glm5-next: give dead indexer slots unique scatter rows (#29745)
The sparse indexer mask is built with a set_rows scatter. Padded pools,
absent sequences and missing tail cells all pointed to the same n_kv
sentinel row, and invisible pools picked by top_k to fill the selection
overlap the tail cells of the token, so several CPU threads wrote the
same element (ThreadSanitizer data race in the sanitize CI).

Allocate the slot mask for both selection paths and route every dead
slot to its own dump row n_kv + slot. Live slots address disjoint cells,
so the scatter indices of a token are unique.
b11302
2026-09-30 17:10:48 +02:00
Daniel Kuts 2149c00f44 ggml/gguf : fix integer overflow (#29384)
* ggml: fix integer overflow guard for zero-element tensors

* ggml: validate number of elements in tensor to prevent integer overflow

* ggml: fix error print
b11301
2026-09-30 17:59:00 +03:00