1. Per row FP8 wrongly fell to NVFP4 conversion, as NVFP4 checked on
dims alone. Also check on dtype for FP8
2. Need to reshape FP8 QKV projections scales in the same way that
weights are reshaped
* hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA
* hex-cpy: various fixes on top of the concat optimizations
Removed CONCAT_DMA_MIN_ROW logic, it was broken with 64-bit DMA.
While it's kinda silly to use DMA for tiny stuff if that tensor gets mapped to an extended buffer the only way to read it is DMA.
Added missing dma_queue_flush() calls.
Added additional guards for conditions we don't support.
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* convert : write Gemma embedding scale for DFlash drafts
A DFlash draft shares the target's token embeddings. Gemma scales them by sqrt(hidden_size) in the forward pass, and the draft config does not state that scale, so the converted draft read unscaled embeddings.
Take the scale from the target config when the draft config has none.
Assisted-by: Claude
* convert : check with get_model_architecture for gemma models
* cuda : route sm70 to the Turing MMVQ nwarps table
Volta (sm_70) has no MMVQ parameter table of its own and falls through
to GENERIC, which launches K-quant batch-1 decode (ncols_dst == 1) at
nwarps=4. sm_70 shares TURING's tuning: the K-quant vec_dot prefers
nwarps=2 there. Route sm_70 to the existing MMVQ_PARAMETERS_TURING
table in both the device and the host table selector.
Measured on one Tesla V100 32GB PCIe (PG500-216, driver 580.178.04,
CUDA 12.0.140) with Qwen3.8-27B Q4_K_M, tg128, interleaved A/B in 6
ABBA blocks with paired per-block deltas: +1.091 t/s = +3.17 %
(t = +49.0, all six per-block deltas positive); perplexity
bit-identical (6.3697 +/- 0.04066 both builds, wiki.test.raw). The
patched build's K-quant mul_mat_vec_q kernels launch at nwarps=2
(cubin EIATTR_MAX_THREADS) while Q4_0/Q8_0 stay at nwarps=4, and the
same measurement on the September master base gave +3.84 % (t = 85).
The tuning originates from the V100-focused fork anyei/llamacpp-v100
(MIT), commit b912d1b1e, which carries a dedicated
MMVQ_PARAMETERS_VOLTA table; a cubin-level comparison confirmed that
routing sm_70 to the existing TURING table is equivalent for the
K-quant batch-1 path this change affects, so this is the minimal
2-line form. https://github.com/anyei/llamacpp-v100/commit/b912d1b1e
Original-patch-by: anyei <angelyoelroblesmercedes@gmail.com>
* Update ggml/src/ggml-cuda/mmvq.cu
---------
Co-authored-by: tkittich <tkittich@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* metal : release temporary private transfer buffers
Assisted-by: OpenAI Codex
* metal : fix order and formatting
---------
Co-authored-by: Niklas Wenzel <dev@nikwen.de>
* CUDA: Handle compute type for NVFP4 on cublass path
Signed-off-by: ynankani <ynankani@nvidia.com>
* Use BF16 compute type for quantized models if HW allows
Signed-off-by: ynankani <ynankani@nvidia.com>
* Set acc prec to bf16 for nvfp4 as it needs atleast bf16 range
Signed-off-by: ynankani <ynankani@nvidia.com>
* Update ggml/src/ggml-cuda/ggml-cuda.cu
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* preserve op_params for per-expert matmul
Signed-off-by: ynankani <ynankani@nvidia.com>
---------
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.
We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)
* BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS
* AOCL-Blas : Add an AOCL-BLAS Quick Start and drop the fixed version path
* AOCL-BLAS doc : Note on ZenDNN
LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space
after every special token and parallel tool calls are separated by
<|end|>. The GPT-OSS handler rejects this output, so add a dedicated
handler, selected by the chat_format=llm-jp-harmony-v1 declaration in
the chat template.
Assisted-by: Claude Fable 5.1
* vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3
The original tokenizer configs for PLaMo-2 and PLaMo-3 have
`add_bos_token: true` and `add_eos_token: false` , but
_set_vocab_plamo() did not write the BOS/EOS metadata. The
PLAMO2 tokenizer path also ignored add_bos/add_eos during
tokenization.
Write the settings from tokenizer_config.json and honor them in
the PLAMO2 tokenization path. GGUFs without these keys keep the
previous behavior.
* Update conversion/base.py
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
The committed docs/ops/CPU.csv is out of sync with the current
test-backend-ops suite: 11 ops with CPU support (COL2IM_1D,
MUL_MAT_HADAMARD, SWIGLU_CLAMP, MUL_MAT_W4A4/W4A8, MUL_MAT_ID_W4A4/W4A8,
DSV4_HC_COMB/PRE/POST, LIGHTNING_INDEXER) are missing entirely, and
many other ops have fewer test cases than the suite generates now.
docs/ops.md (which CI requires to match the CSVs) therefore
understates CPU support.
Regenerated with:
test-backend-ops support -b CPU --output csv > docs/ops/CPU.csv
scripts/create_ops_docs.py
Note: ADD1 now reads unsupported on CPU because ggml_add1 is
GGML_DEPRECATED and the suite no longer generates test cases for it;
the CPU implementation itself is still present.
Assisted-by: Xing
#27941 disabled -sm tensor for qwen4exp because test-llama-archs asserted on the
Meta device once the fixture carried a PLE layer:
GGML_ASSERT(ggml_backend_buffer_is_meta(tensor->buffer)) at ggml-backend-meta.cpp:476.
With host-resident embeddings the PLE gather is a CPU node and hc_init (the REPEAT
that fans the embedding out to the hc streams) was first reached through layer 0's
PLE path, after that gather. ggml_backend_sched_split_graph pass 2 expands a device
assignment upwards only until it meets a CPU node, so the REPEAT stayed on the CPU
and the later reshape of hc_init inside the meta split viewed a host-resident node.
Expanding hc_init right after it is built puts the REPEAT directly before the first
device node, where pass 2 assigns it; the embedding reshape stays in the CPU split
and is copied in as a split input, as in deepseek4.
dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than that (an IQ4_NL row is only guaranteed to be a multiple of QK4_NL). Threads whose 32-value sub-block starts at or past k currently read and write past the end of the row. Skip those sub-blocks; for rows that are a multiple of QK_K the check never fires.
* hex-allreduce: add support for safe scatter mode
* hex-allreduce: pare down excessive comments
* hex-allreduce: re-write to remove register spills
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* llama: preserve original batch order for layer inputs
Assisted-by: Codex
* tests: cover layer-input order across KV layouts
Assisted-by: Codex
* tests: exercise layer-input ordering on CUDA devices
Assisted-by: Codex
* llama: make layer input reordering compatible with tensor split
Copy each microbatch tensor from offset zero and restore original row order after synchronization. Extend the layer-input regression to cover tensor split and repeated reads and decodes.
Assisted-by: Codex
* llama: restore token order for unmasked NextN embeddings
Use the original-token mapping for unmasked NextN rows, including when
layer-input capture is disabled. Keep masked NextN rows on the logits
output mapping and preserve offset-zero tensor copies.
Extend the existing regression to cover NextN alone, combined layer
capture, and masked outputs with repeated decodes and getters.
Validation: all 256 CPU/CUDA/tensor configurations pass. Qwen3.8-27B
Q4_K_M MTP completes MT-Bench at concurrency 16 before and after.
Assisted-by: Codex
* ggml: fix WebGPU reservation and OpenVINO hidden-state capture
Reserve WebGPU vector attention scratch across batch sizes and refresh reservations when NextN capture settings change. Preserve requested OpenVINO outputs, dynamic shapes, sequence counts, and current graph bindings.
Extend existing WebGPU regression coverage and enable strict allocation checks.
Assisted-by: Codex
* llama: defer regression test and backend fixes to follow-ups
Keep this PR focused on restoring token order for layer inputs and unmasked NextN embeddings. Remove the added regression test, OpenVINO and WebGPU changes, and the separate NextN reservation change.
Assisted-by: Codex
* llama: keep n_embd declaration in its original position
Assisted-by: Codex
* llama : pass token count to layer input extraction
Assisted-by: Codex
* llama : name original batch indices batch_idxs
Assisted-by: Codex
* llama : name extracted embedding indices embd_batch_idxs
Assisted-by: Codex
* llama : tag target embedding reordering
Assisted-by: Codex
* llama : tag extraction and name the index capture flag
Assisted-by: Codex
After the device decode, flip causal_attn off, decode n_ubatch/2 then
n_ubatch tokens. Both have the same node count, so a shape that depends
on the flag makes the second reallocate at an unchanged graph size,
which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
* convert: fix LoRA conversion crash for Qwen3.5 V-head reorder
_reorder_v_heads does reshape+permute+reshape to reorder V heads from
grouped to tiled order. LoraTorchTensor.reshape() cannot split its
row dimension (A matrix), so converting Qwen3.5 LoRA adapters that
target out_proj crashes with NotImplementedError.
Fix: detect LoRA tensors and apply the equivalent index permutation
directly — column reorder (dim=last) permutes A's columns, row
reorder (dim=0) permutes B's rows. This is mathematically identical:
(B @ A)[:, perm] == B @ A[:, perm]
(B @ A)[perm, :] == B[perm, :] @ A
Verified: both paths produce exactly zero diff against the full-tensor
reorder on random (rank=32, 4096×4096) matrices.
Fixes#21125
Signed-off-by: Radu Swigler <radu@swigler.com>
* convert: add ty: ignore for hasattr-guarded LoRA call
Assisted-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix comment
* nowrap
---------
Signed-off-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Assisted-by: Claude Opus 4.6 <noreply@anthropic.com>
The sparse indexer mask is built with a set_rows scatter. Padded pools,
absent sequences and missing tail cells all pointed to the same n_kv
sentinel row, and invisible pools picked by top_k to fill the selection
overlap the tail cells of the token, so several CPU threads wrote the
same element (ThreadSanitizer data race in the sanitize CI).
Allocate the slot mask for both selection paths and route every dead
slot to its own dump row n_kv + slot. Live slots address disjoint cells,
so the scatter indices of a token are unique.
* ggml: fix integer overflow guard for zero-element tensors
* ggml: validate number of elements in tensor to prevent integer overflow
* ggml: fix error print