mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-30 18:07:38 -05:00
* hex-concat: reduce pkts in gather/transpose hot loop gather directly into dst buffer, use special instruction for gather sync * hex-concat: use fastdiv replace calls to sw divide with fastpath * hex-concat: optimize DMA-HVX pipeline and add transpose helpers