* BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS
* AOCL-Blas : Add an AOCL-BLAS Quick Start and drop the fixed version path
* AOCL-BLAS doc : Note on ZenDNN
dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than that (an IQ4_NL row is only guaranteed to be a multiple of QK4_NL). Threads whose 32-value sub-block starts at or past k currently read and write past the end of the row. Skip those sub-blocks; for rows that are a multiple of QK_K the check never fires.
* hex-allreduce: add support for safe scatter mode
* hex-allreduce: pare down excessive comments
* hex-allreduce: re-write to remove register spills
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* ggml: fix integer overflow guard for zero-element tensors
* ggml: validate number of elements in tensor to prevent integer overflow
* ggml: fix error print
* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)
* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops
Assisted-by: Claude Opus 5.5
* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build
* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
* cpu: accept BF16 in src1 of mul_mat
ggml_conv_1d_dw builds its im2col in F32 when the kernel is BF16, then
calls ggml_mul_mat(im2col, kernel), which puts F32 in src0 and BF16 in
src1. The CPU backend refused that combination, so it was reported as
unsupported on every backend and never compared against anything.
Widen BF16 into the F32 work buffer, next to the existing packing of F32
into vec_dot_type. This is the arithmetic the Metal mat vec kernel
already uses, both operands promoted to float and accumulated in float,
so the two agree exactly rather than approximately.
Cover it with a conv_1d_dw test over F32, F16 and BF16 kernels, plus
three mul_mat cases with BF16 in src1.
* vulkan: reject BF16 in src1 of mul_mat unless src0 is BF16
supports_op only checked the src1 type for non contiguous tensors, so
a contiguous BF16 src1 was accepted and the pipeline lookup asserted.
The only BF16 src1 path is the BF16 x BF16 multiply, every other src0
type now reports the op as unsupported and the scheduler keeps it on
the CPU.
The BF16 kernel case of the conv_1d_dw test needs the f32 x bf16
mat vec variants of the Metal backend, which land separately.
* openvino: serve GET_ROWS on a weight view from the base Constant
Resolve view_src when collecting weight Constants so a view over a
quantized weight no longer becomes a dynamic typed Parameter, and fold
the row offset of the view into the gather indices instead of slicing
the dequantization subgraph.
* openvino: lift the quantized GET_ROWS view rejection
The supports_op rejection of a quantized src0 view with a nonzero
offset keeps the vs0 GET_ROWS cases of #28253 away from OpenVINO.
The weight view now resolves to the base Constant with the row offset
folded into the gather indices, so the rejection goes away.
The MUSA vendor header never defined __CUDA_ARCH__, so every architecture
test in the shared ggml-cuda sources evaluated to 0. Kernel bodies gated on
the architecture therefore compiled to nothing, for example the q8_0 -> f16
dequantization kernel in convert.cu, whose NO_DEVICE_CODE fallback expands to
an empty body in host code.
Report the newest architecture like the HIP backend does and exclude the
NVIDIA-only features explicitly, as they are not usable on MUSA. Define it
for device passes only: CUB uses defined(__CUDA_ARCH__) to detect device
compilation, which is also how nvcc behaves.
Drop the now-redundant defined(__CUDA_ARCH__) checks in the architecture
comparisons: __CUDA_ARCH__ is undefined in host passes for CUDA and MUSA, and
HIP defines it for every pass, so both forms select the same branch.
* hexagon: add F16 support for activation ops (SILU/GELU/GELU_QUICK/GEGLU/SWIGLU)
Widens ggml_hexagon_supported_activations() to accept F16 (src0/dst/src1
must agree on type), and adds F16 per-thread worker functions in
act-ops.c mirroring the existing F32 workers, backed by new HVX f16
kernels (hvx_sigmoid_f16_aa, hvx_tanh_f16_aa, hvx_mul_mul_f16_aa,
hvx_min_scalar_f16 family).
SILU, GELU, GELU_QUICK, GEGLU, and SWIGLU are verified correct on-device
(QRD8850) via test-backend-ops CPU-diffed correctness tests. SWIGLU_OAI's
F16 path is code-complete and builds clean on host + all 4 DSP arch
variants (v73/v75/v79/v81), but has no F16 test-case coverage in
test-backend-ops and is therefore unverified on-device in this change.
* hex-ops: align macros
* hex-ops: minor formatting
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
GGML_PAD(nbytes, alignment) wraps to 0 when nbytes is within
(alignment - 1) of SIZE_MAX, which silently bypassed the size
overflow guard in gguf_init_from_reader. Reject the tensor before
padding when nbytes + (alignment - 1) would overflow.
Adds a test-gguf handcrafted case (F32, ne = [4, 2^30-1, 2^30+1, 1])
whose ggml_nbytes = 2^64 - 16 lands in the wrap window. Fails on
master, passes with the guard.
* hex-concat: reduce pkts in gather/transpose hot loop
gather directly into dst buffer, use special instruction for gather sync
* hex-concat: use fastdiv
replace calls to sw divide with fastpath
* hex-concat: optimize DMA-HVX pipeline and add transpose helpers
Without CUB (HIP, MUSA) argsort ran the bitonic kernel with one thread
per padded column, so any row above 1024 entries launched an invalid
block configuration. Each thread now owns several columns, every stage
of the network runs all owned columns before the barrier, and the block
is capped at 1024 threads. Shared memory becomes the only bound, which
supports_op checks against the device instead of a fixed 1024.
Rows up to 1024 run the same work as before. Bit-exact with the CUB
path on rows of 2048.
It turns out Intel doesn't particularly like loading F32s one at a
time and we already have the _2aliagned load logic in mul_mat_vec,
so here we use it.
While we do already check all the requirements to load elements 4
at a time across [B]F16 and F32, it turns out [B]F16 loading 4 at a
time is sometimes slower on very specific shapes on Intel BMG.
Loading 4 at a time is a bit faster on F32, but its not material
and I assume might be slower on other platforms.
Note that we also need to validate `a_offset` is 2-aligned in
`mul_mat_vec.comp`, which was missing in the original 2-way-load
patch.
Some selected speedups from `test-backend-ops perf` on a B60.
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=1,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1704 runs - 767.17 us/run - 117.44 MFLOP/run - 153.08 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=1,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 2556 runs - 529.81 us/run - 117.44 MFLOP/run - 221.66 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=2,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1704 runs - 727.13 us/run - 234.88 MFLOP/run - 323.03 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=2,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 2130 runs - 528.84 us/run - 234.88 MFLOP/run - 444.15 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=3,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1704 runs - 702.19 us/run - 352.32 MFLOP/run - 501.74 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=3,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1988 runs - 532.14 us/run - 352.32 MFLOP/run - 662.08 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=4,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1278 runs - 919.50 us/run - 469.76 MFLOP/run - 510.89 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=4,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1917 runs - 543.69 us/run - 469.76 MFLOP/run - 864.03 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=5,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1197 runs - 892.12 us/run - 587.20 MFLOP/run - 658.21 GFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=5,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1881 runs - 575.17 us/run - 587.20 MFLOP/run - 1.02 TFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=8,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1498 runs - 716.40 us/run - 939.52 MFLOP/run - 1.31 TFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=8,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 1819 runs - 576.36 us/run - 939.52 MFLOP/run - 1.63 TFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=512,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 134 runs - 7467.09 us/run - 60.13 GFLOP/run - 8.05 TFLOPS
MUL_MAT(type_a=f32,type_b=f32,m=4096,n=512,k=14336,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0): 134 runs - 7478.12 us/run - 60.13 GFLOP/run - 8.04 TFLOPS
mut_mul_id selected its matmul tile with total token count.
For MoE dispatch grid the true N per workgroup is per-expert rows.
At pp128 on Sarvam 30B that is 6, not 128, so the picker took the l-tile for ~6 live rows.
Most workers in each group had nothing to do.
This wasted time. The slow part was 55% of the whole job.
* fix c++ odr by properly using GGML_COMMON_DECL_CPP
* using actual field rather than macro
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
---------
Co-authored-by: XZiar <xziar@xziar.xziar>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
graph_inputs was populated while splitting the graph, so it only
contained the inputs that are used as srcs of some node. With pipeline
parallelism (n_copies > 1), each graph input contributes n_copies leafs
to graph_copy, so switching between batches that consume different
inputs (e.g. token batches that do not use the embeddings input vs
image batches that do) changed the graph composition. This shifted the
input copies in graph_copy, making the backend ids comparison report
spurious changes and forcing the scheduler to re-reserve. The
re-reserve could then record smaller input sizes (e.g. out_ids with
n_outputs = 0) and abort later on a graph with an unchanged size via
GGML_SCHED_DEBUG_REALLOC.
Collect the inputs after the split instead, from all input leafs of the
graph, so that the graph composition depends only on which inputs
exist, not on which inputs are used.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
* ggml : speed up model loading
A crafted model could hang the server for a very long time, try with:
llama-cli -hf angt/test-gguf-1Mkv -hff model.gguf
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* Avoid empty keys
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* Fix
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
---------
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* tests : use llama_context_ptr in test-recurrent-state-rollback
Replace raw llama_context pointers with llama_context_ptr and drop the
manual llama_free calls and cleanup lambda.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
* tests : run test-recurrent-state-rollback over all dummy models
Add a --models DIR mode that mirrors test-save-load-state: iterate every
dummy model, report PASS/FAIL/SKIP in a table and fail only when a model
fails. Register a single ctest entry with ARGS --models instead of the four
per-model registrations.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
* cont : fix typo
* metal : allow fusing 0-element nodes to keep graph packing shape-independent
The fusion packing in ggml_metal_fusion_max excluded 0-element tensors and
the topk_moe/moe_reduce checks rejected n_tokens == 0, so graphs decoding
batches with no outputs packed differently from the worst-case reserved
graph. The Metal optimizer then reordered the nodes differently and
ggml_gallocr_needs_realloc failed on the layout mismatch, forcing an
unexpected graph re-reserve (caught by GGML_SCHED_DEBUG_REALLOC).
Treat empty tensors like their non-empty counterparts: match them in the
pattern sequence and only reject genuinely malformed shapes. Fused kernels
dispatch zero threadgroups for empty graphs, which is a legal no-op.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
* tests : run test_multi_seq_split_replay as a separate test
test_multi_seq_split_replay was invoked at the end of test_rollback,
so its result was folded into the rollback status and it only ran when
the rollback part passed.
Give it its own test_status return, run both tests independently over
both cache fills via a shared run_tests helper, and report them as
separate rollback / split replay columns in the --models table with
per-test summaries. The exit code fails when either test fails.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
* tests : loosen the split replay nmse bound to 1e-4
test-generate-models seeds its weights from std::random_device, and some
generated lfm2 models drift up to ~1.7e-5 nmse on the split replay due to
rounding noise, tripping the previous 1e-5 bound. Raise the bound to 1e-4
so the random generations stop flaking.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
* tests : reuse run_tests_for_model in single-model mode
The single-model path duplicated the model init and the non-recurrent
check from run_tests_for_model; route it through the shared helper
instead. Model load failures now return FAIL rather than SKIP so that
--model with a broken file still exits non-zero, and the helper loads
with model_only like the --models loop does since the tests create
their own contexts.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
* ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86
* add AVX2 support for masked loading and storing in simd_gemm_ukernel_tail
* ggml-cpu: fix FA softcap handling for padded KV tiles
* metal: support left and circular padding in GGML_OP_PAD
Align Metal with CPU, CUDA and Vulkan: shift the source coordinates by
the left paddings, wrap them around with the same wrap_around when
circular, and read the source through nb00, which also fixes a right
padding of a permuted source. A test case covers it.
Drop the f32_4 kernel: its selection is disabled as slower, and it
fails two pad cases once enabled.
* metal: use a function constant for the circular pad variant
Address review from ggerganov: replace the bool template with FC_PAD,
as FC_upscale_aa does, so the pad kernel is compiled once and
specialized per pipeline.
* vulkan: read the batch stride of an in place src0 from nb[2]
A dim01 contiguous tensor can still be a view whose batches are
strided by more than ne[1] rows, the first rows of a KV cache for
example. Both the mat-vec and the matrix paths read such a tensor in
place but passed ne00*ne01 as the batch stride, so every head past
the first read the wrong rows. The same applies to src1. The stride
now comes from nb[2] whenever the tensor is used in place; the value
is unchanged for a contiguous tensor.
test-backend-ops gets an m_v parameter on test_mul_mat, the number of
rows of a in memory, and two cases at the shapes of a decoder self
attention over a cache.
* vulkan: size the in place A and B ranges by their strided extent
The matrix path bound src0 and src1 to the shader with a range of
elements times type size, which ends before the batches of a strided
view. Pipelines with bounded access read zero past that range, so the
same view that the mat-vec path already handles gave wrong results
on Intel and on NVIDIA without coopmat2. The range now comes from
ggml_nbytes when the tensor is read in place.
* vulkan: address review from jeffbolznv
Bind the in place A and B of the matrix path with ggml_vk_subbuffer,
which spans to the end of the buffer, so a strided view is in range
without computing its extent.
mul_mat_id reads the batch stride of an in place src0 and src1 with
the same helper as mul_mat. test_mul_mat_id gets an m_v parameter,
the number of rows of as in memory, and a case whose experts are
strided by more rows than it uses.
* vulkan: read the batch stride of an in place src0 in mul_mat_vec_id
The single token path of mul_mat_id passed ne00*ne01 as the batch
stride of A, so a strided expert view read the wrong rows. The stride
now comes from ggml_vk_batch_stride like the other three paths, and
src1 follows the same rule.
test_mul_mat_id gets a single token case over the strided view.
* vulkan: address review from jeffbolznv
The batch stride of an in place tensor is taken from nb[2] as
nb[2] / type_size * block_size, which holds when nb[2] is padded and
not a multiple of nb[1]. A test_mul_mat case with a padded batch stride
covers it.
* vulkan: keep the A and B ranges exact in mul_mm
The quantized A loads of mul_mm carry no row bound and rely on the
descriptor range to read zeros past the last row of a partial tile.
Binding A and B up to the end of the buffer let those tiles read the
leftovers of a previous node and hung the NVFP4 mul_mm on NVIDIA
without coopmat2. The range is the strided extent of a tensor read in
place and the staged size otherwise.
The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus
384/640/768/1280 via the Kronecker/Paley construction added separately in
Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through
to the default case and run as a dense GEMM against the materialized
rotation tensor, correct but O(n^2) instead of O(n log n).
fwht_kernel_wide runs one row per work-group instead of per sub-group, so
each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below
the sub-group width still shuffle; those up to the work-group width go
through work-group local memory; the rest stay in registers. Same butterfly
and sign convention as the existing narrow kernel.
ggml's SYCL backend registration (dpct::dev_mgr) unconditionally requires a
GPU-labeled platform to exist and throws before any op-level test can run,
so test-backend-ops could not be exercised on this box (a GPU-less pod) even
via the CPU device. Verified instead with a standalone harness: the same
kernel body run through a real SYCL CPU device (Intel oneAPI DPC++ 2026.1,
OpenCL CPU backend), checked against an independent recursive-doubling
Hadamard reference, cross-validated by first running the existing unmodified
narrow kernel through the identical harness and confirming it passes (rules
out a reference-convention bug before trusting a pass on the new code).
Random-input results for all four widths, single- and multi-row:
N=1024 NT=256 rows=1 max_abs_err=1.7e-07 max_rel_err=4.9e-04 PASS
N=2048 NT=256 rows=1 max_abs_err=1.9e-07 max_rel_err=2.0e-04 PASS
N=4096 NT=256 rows=1 max_abs_err=2.0e-07 max_rel_err=1.4e-04 PASS
N=8192 NT=256 rows=1 max_abs_err=2.5e-07 max_rel_err=3.8e-03 PASS
N=1024 NT=256 rows=7 max_abs_err=2.4e-07 max_rel_err=1.0e-03 PASS
N=2048 NT=256 rows=5 max_abs_err=3.0e-07 max_rel_err=9.4e-04 PASS
N=4096 NT=256 rows=3 max_abs_err=2.7e-07 max_rel_err=1.7e-03 PASS
N=8192 NT=256 rows=2 max_abs_err=2.5e-07 max_rel_err=1.9e-03 PASS
This covers the kernel algorithm itself; it does not exercise the ggml
dispatch/supports_op integration end to end, which needs a real GPU (or a
SYCL GPU plugin) to get past backend registration. test-backend-ops build
is verified: fwht.cpp recompiles with zero warnings as part of ggml-sycl.