Files
llama.cpp/ggml
Ankit Khandelwal 94a0ae3e72 vulkan: MOE aware mat_mul_id tile selection (#29182)
mut_mul_id selected its matmul tile with total token count.
For MoE dispatch grid the true N per workgroup is per-expert rows.
At pp128 on Sarvam 30B that is 6, not 128, so the picker took the l-tile for ~6 live rows.
Most workers in each group had nothing to do.
This wasted time. The slow part was 55% of the whole job.
2026-09-29 20:38:59 +03:00
..
2026-09-23 11:47:24 +03:00
2024-07-13 18:12:39 +02:00
2026-09-24 22:44:05 +03:00