mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-03 11:27:25 -05:00
* metal : add tensor API flash attention kernel for F16 KV * cont : add tensor FA kernels for DK=DV=512 and DK=576, DV=512 * cont : support attention sinks, ALiBi and logit softcap in the tensor FA kernel * cont : add tensor FA kernel for DK=192, DV=128