mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-29 01:17:36 -05:00
* vulkan: read the batch stride of an in place src0 from nb[2] A dim01 contiguous tensor can still be a view whose batches are strided by more than ne[1] rows, the first rows of a KV cache for example. Both the mat-vec and the matrix paths read such a tensor in place but passed ne00*ne01 as the batch stride, so every head past the first read the wrong rows. The same applies to src1. The stride now comes from nb[2] whenever the tensor is used in place; the value is unchanged for a contiguous tensor. test-backend-ops gets an m_v parameter on test_mul_mat, the number of rows of a in memory, and two cases at the shapes of a decoder self attention over a cache. * vulkan: size the in place A and B ranges by their strided extent The matrix path bound src0 and src1 to the shader with a range of elements times type size, which ends before the batches of a strided view. Pipelines with bounded access read zero past that range, so the same view that the mat-vec path already handles gave wrong results on Intel and on NVIDIA without coopmat2. The range now comes from ggml_nbytes when the tensor is read in place. * vulkan: address review from jeffbolznv Bind the in place A and B of the matrix path with ggml_vk_subbuffer, which spans to the end of the buffer, so a strided view is in range without computing its extent. mul_mat_id reads the batch stride of an in place src0 and src1 with the same helper as mul_mat. test_mul_mat_id gets an m_v parameter, the number of rows of as in memory, and a case whose experts are strided by more rows than it uses. * vulkan: read the batch stride of an in place src0 in mul_mat_vec_id The single token path of mul_mat_id passed ne00*ne01 as the batch stride of A, so a strided expert view read the wrong rows. The stride now comes from ggml_vk_batch_stride like the other three paths, and src1 follows the same rule. test_mul_mat_id gets a single token case over the strided view. * vulkan: address review from jeffbolznv The batch stride of an in place tensor is taken from nb[2] as nb[2] / type_size * block_size, which holds when nb[2] is padded and not a multiple of nb[1]. A test_mul_mat case with a padded batch stride covers it. * vulkan: keep the A and B ranges exact in mul_mm The quantized A loads of mul_mm carry no row bound and rely on the descriptor range to read zeros past the last row of a partial tile. Binding A and B up to the end of the buffer let those tiles read the leftovers of a previous node and hung the NVFP4 mul_mm on NVIDIA without coopmat2. The range is the strided extent of a tensor read in place and the staged size otherwise.