Files
llama.cpp/ggml/src
Robert Esclapez e1a1abb787 ggml-cuda: Allow transpose-free gemmv computation (#26171)
When matrix's weights are shaped 1xK is leverage a transpose-free
computation to use mat_mul_vec_f.
2026-07-30 21:39:46 +08:00
..
2026-07-29 15:04:30 +08:00
2026-04-16 17:21:28 +08:00
2026-07-28 16:23:24 +03:00