mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-25 07:27:30 -05:00
the dsv4_hc_pre kernels hardcoded hc = 4 via a constexpr used with simd_shuffle, so the op was rejected by supports_op for any other hc and fell back to CPU. Kimi-K3 uses dsv4_hc_pre with hc equal to the number of banked checkpoints in the cross-layer residual stack, which grows with the layer index. pass n_hc as a function constant (FC_DSV4_HC) with per-n_hc pipeline variants, and loop over it in both pre kernels with direct loads add test-backend-ops cases for hc = 1, 2, 3, 5, 8 and 65, gated and not gated Assisted-by: pi:llama.cpp/Qwen3.8-27B