mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-01 02:17:42 -05:00
* hex-concat: reduce pkts in gather/transpose hot loop gather directly into dst buffer, use special instruction for gather sync * hex-concat: use fastdiv replace calls to sw divide with fastpath * hex-concat: optimize DMA-HVX pipeline and add transpose helpers