mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-30 18:07:38 -05:00
mut_mul_id selected its matmul tile with total token count. For MoE dispatch grid the true N per workgroup is per-expert rows. At pp128 on Sarvam 30B that is 6, not 128, so the picker took the l-tile for ~6 live rows. Most workers in each group had nothing to do. This wasted time. The slow part was 55% of the whole job.