mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-29 09:27:33 -05:00
* models: pad on the left with ggml_pad_ext The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders build a left padding as a right pad followed by a roll, and DFlash2 concatenates a zero filled block in front of the previous tokens. ggml_pad_ext does both in one node now that every backend supports a left padding. The Gemma 4 audio embeddings are bit identical. * models: skip the DFlash2 taps that only read padding A tap at or past block_size shifts every row out of the block, so its term is zero. The loop runs min(kernel_size, block_size) taps.