Commit Graph
5089 Commits
Author SHA1 Message Date
Jhen-Jie Hong fcc2feee2f metal : add metallib build support for xcframework (llama/28163) 2026-09-04 13:39:39 +03:00
anujj 2c486783ab cuda: fuse MoE weighted expert reduction (llama/25952)
* cuda : fuse MoE weighted reduction (mul + view + add)

The MoE combine tail currently writes weighted expert outputs to
global memory before reducing them. That intermediate global-memory
traffic is the main cost. The production baseline generally runs two
physical fused kernels; this path runs one.

This change matches the full expert-weighting plus ordered-reduction
subgraph and replaces it with one weighted-reduction kernel.

Supported graphs:
- unscaled: experts * router_weights
- scaled:   (experts * expert_scale) * router_weights

k = 2..15 is handled by one runtime-k kernel.

Matching is structural: op sequence, shapes, strides, expert views,
and the left-to-right ADD chain. The fused kernel keeps that same
reduction order. Results are not claimed bit-identical; CUDA FP32
contraction can change rounding slightly.

Allocator integration uses add_alloc_dep from the graph-optimizer
API so experts, router weights, and optional expert scales stay live
until the fused destination is written. Memory ranges are rechecked
before the fused kernel runs.

Unrecognized or unsafe graphs are left alone and keep the existing
per-op path. Set GGML_CUDA_MOE_WEIGHTED_REDUCTION=0 to disable the
fusion.

test-backend-ops covers scaled/unscaled, aligned/unaligned, and
representative values across k=2..15, plus a k=16 case that must
stay on the per-op path.

* Pruned the test matrix from 15 to 6

* Addressed the aman and olivers review comments
2026-09-04 13:39:39 +03:00
Titaniumtown f162a19475 Revert "sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…" (#28184)
This reverts commit 1f3d318734c61cf6f3b209726cdd3f9c300a782e.
2026-09-04 13:39:39 +03:00
Jingxin (Philip) Li 408faaafce sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 1280 (llama/28016) 2026-09-04 13:39:38 +03:00
Lukasz Stolcman 5f07f856c5 metal : add fa-vec tuning for M2 Pro (llama/28122)
* metal: add fa-vec tuning for M2 Pro

* metal : update fa-vec tuning for M2 Pro with new dtypes
2026-09-04 13:39:38 +03:00
Jhen-Jie Hong 8cca1a3616 metal : add fa-vec tunings for A18 Pro (MacBook Neo) (llama/28152) 2026-09-04 13:39:38 +03:00
Niklas WenzelandYiChen Lv a245a8f48f metal : fix more leaks due to missing autoreleasepools (llama/27883)
* metal : fix more leaks due to missing autoreleasepools

* metal : rename variable

* metal : fix another missing pool warning

Co-authored-by: YiChen Lv <63285796+forforever73@users.noreply.github.com>

---------

Co-authored-by: YiChen Lv <63285796+forforever73@users.noreply.github.com>
2026-09-04 13:39:38 +03:00
Georgi Gerganov 870db2afa1 metal : add fa-vec tuning for M2 Max (llama/28015)
Rows for M2 Max (30 GPU cores) collected with 'ggml-metal-tuning fa-vec
--dtype f16,q8_0', pasted into fa_vec_tuned_table.

ref: https://github.com/ggml-org/llama.cpp/discussions/27668#discussioncomment-18205786

Assisted-by: pi:llama.cpp/Qwen3.8-27B
2026-09-04 13:39:38 +03:00
Neo Zhang 4f3a2a4b79 sycl : support limit max alloc memory within 2GB for host-pinned memory (llama/27559) 2026-09-04 13:39:38 +03:00
James Francis 5032008bc8 metal: enable Metal 4.0 tensor API on M5+/A19+ (llama/27461)
* metal : request Metal 4.0 language version for the tensor API

* metal : load the tensor API kernels from a separate metallib

* tests : add external-metallib tensor API regression test

* metal : fix metallib build order for the tensor API kernels
2026-09-04 13:39:38 +03:00
Buğra Özgürsoy 8e54c659b5 metal : add fa-vec tunings for M1 Ultra (llama/28088)
* metal : add fa-vec tunings for M1 Ultra

* metal : move M1 Ultra tunings after M1 Max section

* metal : remove duplicate blank line
2026-09-04 13:39:37 +03:00
ynankani f22bb2ea4d CUDA: XOR swizzle flash attn K,V smem fp16 tiles (llama/25635)
* CUDA: XOR swizzle flash attn  K,V smem fp16 tiles

Signed-off-by: ynankani <ynankani@nvidia.com>

* Fix use 64bit generic pointer instead of 32bit shared pointer

Signed-off-by: ynankani <ynankani@nvidia.com>

* fix shared memory race in FA on DGX Spark

* Handle corener case

Signed-off-by: ynankani <ynankani@nvidia.com>

* Add swizzle test cases and gate sync for swizzled path only

Signed-off-by: ynankani <ynankani@nvidia.com>

* gate CUDA PTX

Signed-off-by: ynankani <ynankani@nvidia.com>

* offset calculation specific for swizzle branch

Signed-off-by: ynankani <ynankani@nvidia.com>

* Reafctor code

Signed-off-by: ynankani <ynankani@nvidia.com>

* Refactor FA swizzle ldmatrix if/else into helpers (K row/col, V offset)

Signed-off-by: ynankani <ynankani@nvidia.com>

* rebase and update test case args

Signed-off-by: ynankani <ynankani@nvidia.com>

* Allow swizzle for non-pow2 shapes, for which nbatch_2%32==0

Signed-off-by: ynankani <ynankani@nvidia.com>

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-09-04 13:39:37 +03:00
Georgi Gerganov dbc40efce8 metal : add concat support for quantized types (llama/28116)
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-0731
2026-09-04 13:39:37 +03:00
BartowskiandGeorgi Gerganov 2f608ab4f8 AVX2: Speed up large batch size prompt processing of IQ models (llama/27402)
* Batched gemm for grid IQ quants

Style updates and a bit more performance

Clean up comments

Move code around

Vectorize IQ panel decode, lower threshold for speedup

IQ panel: single-source gather layout, gate bias, vectorize interleave

Add ggml_gemm_iqp_8x8_q8_K_p4 kernel, remove gather buffer

Move IQ panel code out of repack into iqp.cpp, clean up comments

Another comment sweep

* Add myself as iqp.* codeownder

* Remove ggml_cpu_iqp_scratch_offset and ggml_cpu_iqp_src1_conv_size

* Renaming and moving

* The other half of renaming and moving

* Move macros and ggml_cpu_iqp_mul_mat_id_min_batch definition

* Update ggml/src/ggml-cpu/iqp.h

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Add iqp_rows work buffer

* Revert "Add iqp_rows work buffer"

This reverts commit 425542991eee1b01fa3844bf87fc4f205ddbfccb.

* Add NUMA fallback

* Add 10 row batch tests for IQP coverage on all grid IQ types

* Swap assert for return false in support check

* Move IQP mul_mat_id test

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-09-04 13:39:37 +03:00
Georgi Gerganov c6934d0fcf metal : add top-k radix implementation (llama/28073)
Assisted-by: DeepSeek-v4-Flash-0731
2026-09-04 13:39:37 +03:00
Hongqiang Wang 088c603e29 opencl: tune the quant paths for Intel Xe-LP GPUs to improve its TG and PP performance (llama/26438)
* opencl: Q4_K/Q5_K mul_mv N_DST 4->8 on Intel for 2x activation reuse

* opencl: Q4_K mul_mm 8x8 tile fot Intel

* opencl: Q5_K mul_mm 8x8 tile for Intel

* opencl: Q4_K mul_mv N_DST 8->16 for Intel
2026-09-04 13:39:37 +03:00
c648b9a4d0 webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation (llama/28045)
* webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation

* chore : improve code readability

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

---------

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
2026-09-04 13:39:37 +03:00
Jaden_Mach c1be45b890 ROCm: add radix TOP_K for long rows (llama/27466)
* ROCm: add radix TOP_K for long rows
2026-09-04 13:39:36 +03:00
Niklas Wenzel 7614a4c139 metal : add fa-vec tunings for M1 (llama/28078) 2026-09-04 13:39:36 +03:00
ynankani b0f4bc02ed CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token (llama/27621)
* CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were resticted to 1 token

Signed-off-by: ynankani <ynankani@nvidia.com>

* Address review comments

Signed-off-by: ynankani <ynankani@nvidia.com>

* Add SWIGLU_CLAMP case to multi-token moe fusion

Signed-off-by: ynankani <ynankani@nvidia.com>

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-09-04 13:39:36 +03:00
Neo Zhang 76a51e82d7 sycl : Enhance to get the free memory of Intel GPU (llama/27968)
* enhance get mem info by l0 an SYCL API

* remove debug code, format the code

* update SYCL.md for GGML_SYCL_GET_MEM_API
2026-09-04 13:39:36 +03:00
Simon Teixidor 6ce7b89952 vulkan: tune mat-vec rows for batched inference on Strix Halo (llama/27909)
* vulkan: RDNA3 static mat-vec rows above four columns

On RDNA3 above four columns a static 4 rows for all types benches faster than
the default.

* vulkan: RDNA3 static mat-vec-id rows

mul_mat_vec_id has no column dimension to switch on. On my Strix Halo machine,
a static 4 is faster here than the defaults across types and batch sizes.
2026-09-04 13:39:36 +03:00
fairydreamingandStanisław Szymczyk 96dddd87f2 ggml : add MUL_MAT to the list of ops that may need additional memory (for WebGPU) (llama/28071)
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
2026-09-04 13:39:36 +03:00
Ruben Ortlam db00b0196b vulkan: top_k radix select for k >= 1024 for Qwen 3.8 Flash Next (llama/28032)
* vulkan: add top-k radix sort shader for k >= 1024

* add Qwen 3.8 Flash Next top-k tests

* add top-k qsa fusion

* clean up code
2026-09-04 13:39:36 +03:00
Shenghan Yang 01ebd225a7 hexagon: fix CPY fence bug (llama/28033) 2026-09-04 13:39:35 +03:00
codemonkey e5c96ca45d metal : add remaining Q4_1/Q5_0/Q5_1 fa-vec tunings for M2 (llama/28017) 2026-09-04 13:39:35 +03:00
hmirinandGeorgi Gerganov 4089fa628a rpc: avoid serializing buffers from other servers (llama/26500)
* rpc: avoid serializing buffers from other servers

Only include remote buffer pointers when the buffer belongs to the RPC dispatcher receiving the graph. Add a two-server regression test for cross-server tensor serialization.

Assisted-by: Codex

* cont : add ref

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-09-04 13:39:35 +03:00
Georgi Gerganov 749683d30f ggml : fix ggml_backend_buft_get_alloc_size() guard (llama/28038) 2026-09-04 13:39:35 +03:00
Aman Gupta e9583f075a ggml: add SWIGLU_CLAMP (llama/27930)
* ggml: add SWIGLU_CLAMP

* add vulkan shader
2026-09-04 13:39:35 +03:00
Pascal e900a732c8 CUDA: use the fast mm_ids_helper path for any n_expert_used (llama/27978)
The optimized path grouped warp lanes by token and required
warp_size % n_expert_used == 0, with a single hardcoded exception
padding 6 up to 8. Every other count fell back to the generic path,
which walks the tokens one at a time with a warp reduction per token,
for each of the n_expert blocks.

The lane group only has to divide the warp, and the loop body already
guards the padded lanes with iex < n_expert_used, so the padding
generalizes to the next power of two. The 6 -> 8 case and every count
already dispatched keep the exact same padding as before.

n_expert_used = 10 now reaches the fast path. Measured on
Qwen3.8-Flash-Next (512 experts, 10 used) at 55k context on an
RTX PRO 6000, warm runs with the first one discarded:

  prompt processing   2334 -> 2600 t/s

Token generation is unaffected, since a single token leaves nothing to
walk. Other expert counts reach the fast path by adding their case to
the dispatch.
2026-09-04 13:39:35 +03:00
itterative 35d9e2237e hip: tune rdna 3 mmq config (llama/26284) 2026-09-04 13:39:34 +03:00
LunalFresh e5c9e3e3e9 hip : optimize Q2_0 dot-product path for gfx1201 (llama/26753)
* hip/gfx1201: optimize q2_0 vec_dot_q2_0_q8_1 with native amdgcn perm

* Broadened HIP's Q2_0 perm optimization

* Remove redundant HIP perm availability guard

* Optimize HIP Q2_0 MMQ unpack with native perm

* cuda: label HIP preprocessor guard

* cuda: label HIP preprocessor guard

* Restore MMQ tile index handling
2026-09-04 13:39:34 +03:00
Georgi Gerganov 4b2243a6c2 ggml : add ggml_backend_op_alloc_size_may_expand, use it in RPC (llama/27960)
some backends (Metal, SYCL, WebGPU) require additional memory for
fleeting data for certain ops, which is reflected in their
get_alloc_size implementations.

add ggml_backend_op_alloc_size_may_expand() to the backend utils,
listing these ops, and assert in ggml_backend_buft_get_alloc_size
that a backend expanding the alloc size of a compute op only does so
for ops listed in the helper.

use the helper in the RPC backend to decide whether to query the
remote server for the actual alloc size, instead of a hardcoded list.

Assisted-by: pi:llama.cpp/Qwen3.8-27B
2026-09-04 13:39:34 +03:00
Ryan C 43acf3d6e8 rpc: fix apple rdma error spew on teardown (llama/27908) 2026-09-04 13:39:34 +03:00
Nils Gladitz 1e0f382573 metal: add fa-vec tunings for M3 Ultra (llama/27999) 2026-09-04 13:39:34 +03:00
Daya Adianto b66593ef1b metal : Add fa-vec tuning for M3 Pro (llama/27963)
Related issue: #27668
2026-09-04 13:39:34 +03:00
Ryan C 5e494599a4 rpc : fix pre-rdma macOS versions (llama/27815) 2026-09-04 13:39:34 +03:00
3ad8b9b217 hexagon: support for device discovery and create sessions on demand (llama/27785)
* hex-devices: add support for lazy session allocation and cleanup dev interfaces

Co-authored-by: Marco Colombo <mcolombo@qti.qualcomm.com>

* hex-devices: support for runtime discovery of available NPU cores

Co-authored-by: Alexander Lu <alexlu@qti.qualcomm.com>
Co-authored-by: Ehsan Bateni <ebateni@qti.qualcomm.com>

* hex-devices: reject non-existing devices early during init

---------

Co-authored-by: Marco Colombo <mcolombo@qti.qualcomm.com>
Co-authored-by: Alexander Lu <alexlu@qti.qualcomm.com>
Co-authored-by: Ehsan Bateni <ebateni@qti.qualcomm.com>
2026-09-04 13:39:33 +03:00
Titaniumtown 3d4e0e9858 sycl: split long rows in TOP_K instead of one work-group per row (llama/27847) 2026-09-04 13:39:33 +03:00
QuintinShaw c68f20555a metal : fix null-pipeline crash for F16 src1 mul_mat/mul_mat_id (llama/25648)
* metal : fail closed on mul_mat shapes with missing F16 kernels

* metal : abort on nil pipeline in encoder_set_pipeline

* metal : address review comments

* metal : share mul_mat mm dispatch with supports_op
2026-09-04 13:39:33 +03:00
Aman Gupta c969c68b54 ggml: allow passing alloc dependencies in graph_optimize (llama/27301)
* ggml: allow passing alloc dependencies in graph_optimize

* add alloc dep tests

* add TODO about using flat array
2026-09-04 13:39:33 +03:00
codemonkey b33bbc5c4e metal : add fa-vec tunings for M2 (llama/27940) 2026-09-04 13:39:33 +03:00
Hongqiang WangandLi He 2a11026cfe opencl: use a better matmul path on two Adreno GPU generations (llama/27640)
* opencl: default the Adreno xmem F16xF32 GEMM on for X2E

kernel_mul_mm_f16_f32_l4_lm is the slowest matmul this backend has on Adreno: on
the X2-90 it runs the gpt-oss-20b attention projections at roughly a quarter of
what the tuned dense q4_0 GEMM reaches on the same device. That matters for any
model whose non-expert weights stay f16 -- the stock gpt-oss-20b release is
exactly that, and its prefill spends 40.8% of GPU time in that one kernel. The
xmem route already existed but was left opt-in, so nobody hit it.

Worth about 25% prefill on gpt-oss-20b on an Adreno X2-90. Gated to X2E: the
Adreno 840 measures neutral. Decode is untouched -- the dispatch gate needs
N >= 16. It is worth nothing on the q8attn variant, whose attention weights
already take the dp4a dense GEMM.

The env var was presence-tested before, so =0 previously enabled it; it is now
atoi()'d. MUL_MAT 963 OK / 0 FAIL on both arms.

* opencl: bypass the tiled f32 GEMM on the Adreno A7X

The A7X (E031.41) compiler executes kernel_mul_mm_f32_f32_l4_lm at roughly a
tenth of what the same silicon reaches in its own f16 and q4_K kernels. It
allocates 488 B/WI of private memory against 304 for the same source on the
following generation, i.e. the older register allocator spills in the K-loop.
Models with per-layer F32 projection pairs kept F32 by quantization policy land
on this kernel twice per layer, and it dominates their prefill on that part.

Route batched f32xf32 (ne11 > 8) around the tiled path on the A7X and let it
fall through to the per-row f32 kernel, which that compiler handles fine; small
batches keep the tiled path. Weights stay GPU-resident, so decode placement is
untouched -- declining the op in supports_op instead was measured first and
rejected, because the per-layer CPU round-trips cost more decode than the
prefill it gained.

Worth about 9% prefill on gemma-3n-E4B on an Adreno 740, with MUL_MAT counts
identical on and off. No other generation is affected. Override with
GGML_OPENCL_A7X_F32_LM_BYPASS=0.

* opencl: enable xmem GEMM for adreno by default

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
2026-09-04 13:39:33 +03:00
Georgi Gerganov 285f1ffd99 metal : assert shared memory padding (llama/27951)
* metal : assert shared memory padding

* cont : add ref
2026-09-04 13:39:33 +03:00
Niklas Wenzel 308fa4f8a2 metal : add remaining fa-vec tunings for M4 Pro (llama/27915) 2026-09-04 13:39:32 +03:00
Nick Farrell 325c8d16c1 sycl: make --fit respect --fit-target better (llama/27629)
improve the --fit algorithm to take into account the actual peak
required VRAM for a given context size on a SYCL backend.

This includes both properly accounting for how much VRAM is required
when the allocated context is fully used (which makes the reported
context drop below what it did before, but stop it OOMing) as well
as preventing some overly-conservative calculations which meant too much
VRAM was being reserved.

Tested on a Arc b70 with unsloth's qwen3.8 (Q4_K_XL), able to get 262144 context,
fully usable, with q8_0 KV and MTP and 4k ubatch size using --fit-target 1
2026-09-04 13:39:32 +03:00
Jeff Bolz e644752070 vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs (llama/27526)
* vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs

* remove one more fastdiv
2026-09-04 13:39:32 +03:00
Jhen-Jie Hong 590fe18902 metal : add fa-vec tunings for M1 Max (llama/27932) 2026-09-04 13:39:32 +03:00
Jeff Bolz d501a0a3e8 vulkan: Change mul_mat_id to pad K rather than N (llama/27925)
The N padding is needed for mul_mat, but not mul_mat_id. For mul_mat_id,
we indirect the row index through a shared memory lookup table which avoids
any OOB row coordinate. But that callback doesn't bounds check K, so we
actually need K padding instead.
2026-09-04 13:39:32 +03:00
Eric A StaleeandJeff Bolz 4c38040fd3 vulkan: fix missing view-alias dependencies in ggml_vk_graph_optimize (llama/27812)
* vulkan: fix missing view-alias dependencies in ggml_vk_graph_optimize

is_src_of doesn't treat two views of one tensor as dependent, so the optimizer reorders nodes across aliased reads and writes.

Result: silently wrong tokens under greedy decoding, different output on every server start, and invalid speculative-decoding acceptance, with nothing logged.

Hits Qwen3.8's recurrent state (and any model with view-aliased state) on AMD and NVIDIA Vulkan.  CUDA is clean.

Compare view_src bases on both sides.

Fixes #27805

* vulkan: don't treat view/no-op nodes as aliasing dependencies

Nodes whose op is NONE, RESHAPE, TRANSPOSE, VIEW or PERMUTE execute nothing, so aliasing through them is not a real dependency. The previous base comparison matched them anyway, which only costs the optimizer reordering freedom.

Co-authored-by: Jeff Bolz <jbolz@nvidia.com>

* vulkan: make the lambda parameter const and capture is_empty in is_src_of

Code will not compile without these changes.
is_src_of has an empty capture list, so is_empty was not visible inside it, and is_empty took a non-const pointer, while is_src_of receives const ones. Other call sites pass non-const pointers, which still convert as usual.

---------

Co-authored-by: Jeff Bolz <jbolz@nvidia.com>
2026-09-04 13:39:32 +03:00