mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-02 10:57:33 -05:00
* cuda : route sm70 to the Turing MMVQ nwarps table Volta (sm_70) has no MMVQ parameter table of its own and falls through to GENERIC, which launches K-quant batch-1 decode (ncols_dst == 1) at nwarps=4. sm_70 shares TURING's tuning: the K-quant vec_dot prefers nwarps=2 there. Route sm_70 to the existing MMVQ_PARAMETERS_TURING table in both the device and the host table selector. Measured on one Tesla V100 32GB PCIe (PG500-216, driver 580.178.04, CUDA 12.0.140) with Qwen3.8-27B Q4_K_M, tg128, interleaved A/B in 6 ABBA blocks with paired per-block deltas: +1.091 t/s = +3.17 % (t = +49.0, all six per-block deltas positive); perplexity bit-identical (6.3697 +/- 0.04066 both builds, wiki.test.raw). The patched build's K-quant mul_mat_vec_q kernels launch at nwarps=2 (cubin EIATTR_MAX_THREADS) while Q4_0/Q8_0 stay at nwarps=4, and the same measurement on the September master base gave +3.84 % (t = 85). The tuning originates from the V100-focused fork anyei/llamacpp-v100 (MIT), commit b912d1b1e, which carries a dedicated MMVQ_PARAMETERS_VOLTA table; a cubin-level comparison confirmed that routing sm_70 to the existing TURING table is equivalent for the K-quant batch-1 path this change affects, so this is the minimal 2-line form. https://github.com/anyei/llamacpp-v100/commit/b912d1b1e Original-patch-by: anyei <angelyoelroblesmercedes@gmail.com> * Update ggml/src/ggml-cuda/mmvq.cu --------- Co-authored-by: tkittich <tkittich@gmail.com> Co-authored-by: Johannes Gäßler <johannesg@5d6.de>