Files
llama.cpp/ggml
Masashi Yoshimura 8f4646a63e ggml-webgpu: improve flash_attn_vec for quantized KV at long contexts (#25956)
* improve fa of quantized kv cache

* Fix some bugs and some comments.

* fix v type check and some comments

* Fix build error caused by rebasing

* editorconfig checking pass
2026-07-31 09:08:40 +03:00
..
2026-07-29 15:04:30 +08:00
2024-07-13 18:12:39 +02:00