Files
Georgi Gerganov 5b59b83f4e metal : add MoE and SSM_CONV fusion optimizations (#28948)
* metal : add top-k MoE fusion

Adds a Metal fusion for SOFT_MAX + ARGSORT + GET_ROWS with optional
routing-weight normalization and scale, matching the top-k MoE fusion
available in the CUDA and Vulkan backends. The fused kernel writes the
selected expert ids and routing weights directly, eliding the separate
softmax, argsort, get-rows, sum-rows, clamp, div and scale kernels.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : add MoE weighted reduction fusion

Fuses MUL(experts, weights) plus the expert VIEW/ADD chain into one kernel
that computes the weighted sum directly. The graph_optimize hook keeps the
expert and weight buffers alive until the fused output so the allocator cannot
reuse them while the kernel is still reading them.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* tests : expose MoE weighted reduction in fusion baseline

Use 2 experts per token in the generated MoE test models so the Metal
MoE weighted reduction fusion (MUL + ADD) is exercised by test-fusion.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : fuse RMS_NORM + SCALE

Adds NORM/RMS_NORM + SCALE fusion to the Metal backend by reusing the
norm+mul kernel with a scalar scale flag. Adds test coverage for both
NORM+SCALE and RMS_NORM+SCALE and regenerates the fusion baseline.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : use function constant for RMS_NORM + SCALE

Replaces the runtime use_scale karg with a Metal function constant. The
norm+mul kernel is compiled with FC_norm_use_scale=false for MUL fusion and
FC_norm_use_scale=true for SCALE fusion, so the fused kernel has no runtime
branch.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : use function constant for top-k MoE with_norm

Replaces the runtime with_norm karg with a Metal function constant. The
top-k MoE kernel is compiled separately for the normalized and non-normalized
routing variants, removing the runtime branch.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : rename moe_weighted_reduction suffix to moe_reduce

Shortens the MoE weighted-reduction fusion identifiers, kernel, pipeline,
matcher, args struct, and test op name from moe_weighted_reduction to
moe_reduce.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : add MUL_MAT + UNARY and MUL_MAT + ADD + UNARY fusion

Adds dense mat-vec activation fusion for sigmoid/silu and bias+softplus.
The mat-vec kernels apply the activation/bias epilogue via function
constants, avoiding the separate unary/add passes.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : revert MUL_MAT + UNARY and MUL_MAT + ADD + UNARY fusion

The mat-vec activation fusion regressed decode throughput on Qwen3.6-35B-A3B
by ~8% (tg32 81.5 vs 88.5 t/s). The regression is caused by loss of
concurrency: the standalone unary kernels previously overlapped with other
mat-vec work, while fusing the activation into the mat-vec kernel serializes
it on the critical path.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : add SSM_CONV + UNARY (silu) fusion

The SSM_CONV kernels apply silu directly via a function constant, eliding
the separate unary pass. Regenerates the fusion baseline.

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : address fusion review comments

- Fix declaration/table alignment
- Rename top-k MoE kargs fields to val_clamp / val_scale
- Move moe-reduce alloc-deps handling into a general fusion helper
- Remove the public moe-reduce matcher API

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : fix unused parameter in top-k MoE fusion check

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : guard SSM_CONV fusion lookup behind use_fusion

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : track all fused outputs in graph reorder

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : keep top-k MoE logits alive until fused output

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : refactor alloc deps to pattern-driven approach

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : check fused kernel destination in concurrency tracking

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* meta : forward graph_optimize to underlying backends

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : use vector for fusion table

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* meta : keep graph_optimize unimplemented

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* parallel : fix non-deterministic prompt selection

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* parallel : support dummy models and add global logits run hash

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : sync cross-device copies with destination completion event

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : avoid const_cast in fusion alloc deps

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : skip fusions with aliased sources

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : hide fusion pattern definition

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : use vector fusion op sequences

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : drop redundant struct keywords

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : add alloc deps comment separator

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : generalize fusion output memory ranges

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : rename fusion out_offsets to outs

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : avoid dst vector in memory range check

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : optimize fusion matching and multi-output handling

- use pointer arithmetic for fusion info count lookup
- avoid heap allocations in top-k MoE and MoE reduce pattern matchers
- use fusion outs for multi-output subgraph checks

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* Revert "parallel : support dummy models and add global logits run hash"

This reverts commit 57c7caf941c1b43c270fd5009c9f175063522e96.

* fusion : update MTL.csv

* metal : unroll constant loops in top-k MoE kernel

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : use function constants for top-k MoE n_expert and top_k

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : rename fusion kargs to scale and clamp

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* metal : use function constants for moe_reduce and ssm_conv

Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp

* fusion : update MTL.csv
2026-09-19 13:14:44 +03:00

267 lines
19 KiB
CSV

# test-fusion baseline for device MTL
# arch ,moe ,mode ,label , count
afmoe ,1 ,any ,MUL+ADD , 2
afmoe ,1 ,any ,RMS_NORM+MUL , 10
afmoe ,1 ,any ,RMS_NORM+MUL+ADD , 3
arcee ,0 ,any ,RMS_NORM+MUL , 5
arctic ,0 ,any ,MUL+ADD , 4
arctic ,0 ,any ,RMS_NORM+MUL , 7
arctic ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
baichuan ,0 ,any ,RMS_NORM+MUL , 5
bailingmoe ,1 ,any ,ADD+ADD , 2
bailingmoe ,1 ,any ,MUL+ADD , 4
bailingmoe ,1 ,any ,RMS_NORM+MUL , 5
bailingmoe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS , 2
bailingmoe2 ,1 ,any ,ADD+ADD , 1
bailingmoe2 ,1 ,any ,MUL+ADD , 2
bailingmoe2 ,1 ,any ,RMS_NORM+MUL , 9
bailingmoe3 ,1 ,any ,ADD+ADD , 1
bailingmoe3 ,1 ,any ,GATED_DELTA_NET+CPY , 1
bailingmoe3 ,1 ,any ,MUL+ADD , 2
bailingmoe3 ,1 ,any ,RMS_NORM+MUL , 8
bailingmoe3 ,1 ,any ,RMS_NORM+SCALE , 2
bloom ,0 ,any ,NORM+MUL+ADD , 6
chatglm ,0 ,any ,RMS_NORM+MUL , 5
codeshell ,0 ,any ,NORM+MUL+ADD , 5
cogvlm ,0 ,any ,RMS_NORM+MUL , 5
cohere2 ,0 ,any ,ADD+ADD , 2
cohere2 ,0 ,any ,NORM+MUL , 3
cohere2moe ,1 ,any ,ADD+ADD , 2
cohere2moe ,1 ,any ,MUL+ADD , 2
cohere2moe ,1 ,any ,RMS_NORM+MUL , 3
command-r ,0 ,any ,NORM+MUL , 3
dbrx ,0 ,any ,MUL+ADD , 4
dbrx ,0 ,any ,NORM+MUL , 5
dbrx ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
deci ,0 ,any ,RMS_NORM+MUL , 5
deepseek ,0 ,any ,ADD+ADD , 1
deepseek ,0 ,any ,MUL+ADD , 2
deepseek ,0 ,any ,RMS_NORM+MUL , 5
deepseek ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS , 1
deepseek2 ,0 ,any ,ADD+ADD , 1
deepseek2 ,0 ,any ,MUL+ADD , 2
deepseek2 ,0 ,any ,RMS_NORM+MUL , 9
deepseek32 ,0 ,any ,ADD+ADD , 1
deepseek32 ,0 ,any ,MUL+ADD , 2
deepseek32 ,0 ,any ,NORM+MUL+ADD , 2
deepseek32 ,0 ,any ,RMS_NORM+MUL , 9
deepseek4 ,0 ,any ,MUL+ADD , 8
deepseek4 ,0 ,any ,RMS_NORM+MUL , 20
dots1 ,0 ,any ,ADD+ADD , 1
dots1 ,0 ,any ,MUL+ADD , 2
dots1 ,0 ,any ,RMS_NORM+MUL , 9
dots3note ,0 ,any ,ADD+ADD , 1
dots3note ,0 ,any ,MUL+ADD , 2
dots3note ,0 ,any ,NORM+MUL+ADD , 1
dots3note ,0 ,any ,RMS_NORM+MUL , 11
dream ,0 ,any ,RMS_NORM+MUL , 5
ernie4_5-moe ,1 ,any ,ADD+ADD , 1
ernie4_5-moe ,1 ,any ,MUL+ADD , 2
ernie4_5-moe ,1 ,any ,RMS_NORM+MUL , 5
ernie4_5 ,0 ,any ,RMS_NORM+MUL , 5
exaone ,0 ,any ,RMS_NORM+MUL , 5
exaone-moe ,1 ,any ,ADD+ADD , 1
exaone-moe ,1 ,any ,MUL+ADD , 2
exaone-moe ,1 ,any ,RMS_NORM+MUL , 9
exaone4 ,0 ,any ,RMS_NORM+MUL , 5
exaone4 ,0 ,any ,RMS_NORM+MUL+ADD , 4
falcon ,0 ,any ,ADD+ADD , 2
falcon ,0 ,any ,NORM+MUL+ADD , 5
falcon-h1 ,0 ,any ,ADD+ADD , 2
falcon-h1 ,0 ,any ,RMS_NORM+MUL , 9
gemma ,0 ,any ,RMS_NORM+MUL , 5
gemma2 ,0 ,any ,RMS_NORM+MUL , 5
gemma2 ,0 ,any ,RMS_NORM+MUL+ADD , 4
gemma3 ,0 ,any ,RMS_NORM+MUL , 9
gemma3 ,0 ,any ,RMS_NORM+MUL+ADD , 4
glm-dsa ,0 ,any ,ADD+ADD , 1
glm-dsa ,0 ,any ,MUL+ADD , 2
glm-dsa ,0 ,any ,NORM+MUL+ADD , 2
glm-dsa ,0 ,any ,RMS_NORM+MUL , 9
glm4 ,0 ,any ,RMS_NORM+MUL , 5
glm4 ,0 ,any ,RMS_NORM+MUL+ADD , 4
glm4moe ,1 ,any ,ADD+ADD , 1
glm4moe ,1 ,any ,MUL+ADD , 2
glm4moe ,1 ,any ,RMS_NORM+MUL , 9
gpt-oss ,0 ,any ,MUL+ADD , 4
gpt-oss ,0 ,any ,RMS_NORM+MUL , 5
gpt2 ,0 ,any ,NORM+MUL+ADD , 5
gptneox ,0 ,any ,NORM+MUL+ADD , 5
granite ,0 ,any ,RMS_NORM+MUL , 5
granite ,0 ,any ,MUL+ADD , 4
granite ,0 ,any ,RMS_NORM+MUL , 5
granite ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
granite_swa ,0 ,any ,RMS_NORM+MUL , 5
granitehybrid ,0 ,any ,RMS_NORM+MUL , 6
granitemoe ,1 ,any ,RMS_NORM+MUL , 5
granitemoe ,1 ,any ,MUL+ADD , 4
granitemoe ,1 ,any ,RMS_NORM+MUL , 5
granitemoe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
grok ,0 ,any ,MUL+ADD , 4
grok ,0 ,any ,RMS_NORM+MUL , 5
grok ,0 ,any ,RMS_NORM+MUL+ADD , 4
grok ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
grovemoe ,1 ,any ,ADD+ADD , 2
grovemoe ,1 ,any ,MUL+ADD , 8
grovemoe ,1 ,any ,RMS_NORM+MUL , 9
hunyuan-dense ,0 ,any ,RMS_NORM+MUL , 9
hunyuan-moe ,1 ,any ,ADD+ADD , 2
hunyuan-moe ,1 ,any ,MUL+ADD , 4
hunyuan-moe ,1 ,any ,RMS_NORM+MUL , 9
hunyuan-moe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
hunyuan_vl ,0 ,any ,RMS_NORM+MUL , 9
hy_v3 ,0 ,any ,ADD+ADD , 2
hy_v3 ,0 ,any ,MUL+ADD , 4
hy_v3 ,0 ,any ,RMS_NORM+MUL , 9
hy_v4 ,0 ,any ,MUL+ADD , 2
hy_v4 ,0 ,any ,NORM+MUL+ADD , 1
hy_v4 ,0 ,any ,RMS_NORM+MUL , 9
internlm2 ,0 ,any ,RMS_NORM+MUL , 5
jais ,0 ,any ,NORM+MUL+ADD , 5
jais2 ,0 ,any ,NORM+MUL+ADD , 5
jamba ,0 ,any ,RMS_NORM+MUL , 8
kimi-k3 ,0 ,any ,GATED_DELTA_NET+CPY , 1
kimi-k3 ,0 ,any ,MUL+ADD , 2
kimi-k3 ,0 ,any ,RMS_NORM+MUL , 17
kimi-k3 ,0 ,any ,RMS_NORM+SCALE , 2
kimi-linear ,0 ,any ,ADD+ADD , 1
kimi-linear ,0 ,any ,GATED_DELTA_NET+CPY , 1
kimi-linear ,0 ,any ,MUL+ADD , 2
kimi-linear ,0 ,any ,RMS_NORM+MUL , 7
kimi-linear ,0 ,any ,RMS_NORM+SCALE , 2
laguna ,0 ,any ,ADD+ADD , 1
laguna ,0 ,any ,MUL+ADD , 2
laguna ,0 ,any ,RMS_NORM+MUL , 9
lfm2 ,0 ,any ,RMS_NORM+MUL , 7
lfm2moe ,1 ,any ,MUL+ADD , 2
lfm2moe ,1 ,any ,RMS_NORM+MUL , 7
llada ,0 ,any ,RMS_NORM+MUL , 5
llada-moe ,1 ,any ,MUL+ADD , 4
llada-moe ,1 ,any ,RMS_NORM+MUL , 9
llada-moe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS , 2
llama ,0 ,any ,RMS_NORM+MUL , 5
llama ,0 ,any ,MUL+ADD , 4
llama ,0 ,any ,RMS_NORM+MUL , 5
llama ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
llama4 ,0 ,any ,ADD+ADD , 2
llama4 ,0 ,any ,RMS_NORM+MUL , 9
maincoder ,0 ,any ,RMS_NORM+MUL , 9
mamba ,0 ,any ,RMS_NORM+MUL , 3
mamba2 ,0 ,any ,RMS_NORM+MUL , 5
maple ,0 ,any ,MUL+ADD , 4
maple ,0 ,any ,RMS_NORM+MUL , 9
maple ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
mellum ,0 ,any ,MUL+ADD , 4
mellum ,0 ,any ,RMS_NORM+MUL , 9
mellum ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
mimo2 ,0 ,any ,MUL+ADD , 4
mimo2 ,0 ,any ,RMS_NORM+MUL , 5
minicpm ,0 ,any ,RMS_NORM+MUL , 5
minicpm ,0 ,any ,MUL+ADD , 4
minicpm ,0 ,any ,RMS_NORM+MUL , 5
minicpm ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
minicpm3 ,0 ,any ,RMS_NORM+MUL , 9
minimax-01 ,0 ,any ,MUL+ADD , 4
minimax-01 ,0 ,any ,RMS_NORM+MUL , 6
minimax-01 ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
minimax-m2 ,0 ,any ,MUL+ADD , 4
minimax-m2 ,0 ,any ,RMS_NORM+MUL , 9
minimax-m3 ,0 ,any ,ADD+ADD , 1
minimax-m3 ,0 ,any ,MUL+ADD , 2
minimax-m3 ,0 ,any ,RMS_NORM+MUL , 11
mistral3 ,0 ,any ,RMS_NORM+MUL , 5
mistral3 ,0 ,any ,MUL+ADD , 4
mistral3 ,0 ,any ,RMS_NORM+MUL , 5
mistral3 ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
mistral4 ,0 ,any ,ADD+ADD , 1
mistral4 ,0 ,any ,MUL+ADD , 2
mistral4 ,0 ,any ,RMS_NORM+MUL , 9
mpt ,0 ,any ,NORM+MUL+ADD , 5
muse-glimmer ,0 ,any ,RMS_NORM+MUL , 10
muse-glimmer ,0 ,any ,RMS_NORM+MUL+ADD , 3
nanbeige ,0 ,any ,RMS_NORM+MUL , 5
nemotron ,0 ,any ,NORM+MUL+ADD , 5
nemotron_h ,0 ,any ,RMS_NORM+MUL , 5
nemotron_h_moe ,1 ,any ,RMS_NORM+MUL , 5
olmo2 ,0 ,any ,RMS_NORM+MUL , 5
olmo2 ,0 ,any ,RMS_NORM+MUL+ADD , 4
olmoe ,1 ,any ,MUL+ADD , 4
olmoe ,1 ,any ,RMS_NORM+MUL , 9
olmoe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS , 2
openelm ,0 ,any ,RMS_NORM+MUL , 9
orion ,0 ,any ,NORM+MUL+ADD , 5
paddleocr ,0 ,any ,RMS_NORM+MUL , 5
pangu-embedded ,0 ,any ,RMS_NORM+MUL , 5
phi2 ,0 ,any ,ADD+ADD , 2
phi2 ,0 ,any ,NORM+MUL+ADD , 3
phi3 ,0 ,any ,RMS_NORM+MUL , 5
phimoe ,1 ,any ,MUL+ADD , 4
phimoe ,1 ,any ,RMS_NORM+MUL+ADD , 5
phimoe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
plamo ,0 ,any ,ADD+ADD , 2
plamo ,0 ,any ,RMS_NORM+MUL , 3
plamo2 ,0 ,any ,RMS_NORM+MUL , 10
plamo2 ,0 ,any ,RMS_NORM+MUL+ADD , 4
plamo2 ,0 ,any ,SSM_CONV+UNARY , 1
plamo3 ,0 ,any ,RMS_NORM+MUL , 9
plamo3 ,0 ,any ,RMS_NORM+MUL+ADD , 4
pockettts ,0 ,any ,NORM+MUL+ADD , 5
qwen ,0 ,any ,RMS_NORM+MUL , 5
qwen2 ,0 ,any ,RMS_NORM+MUL , 5
qwen2moe ,1 ,any ,ADD+ADD , 2
qwen2moe ,1 ,any ,MUL+ADD , 4
qwen2moe ,1 ,any ,RMS_NORM+MUL , 5
qwen2moe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS , 2
qwen2vl ,0 ,any ,RMS_NORM+MUL , 5
qwen3 ,0 ,any ,RMS_NORM+MUL , 9
qwen35 ,0 ,any ,GATED_DELTA_NET+CPY , 1
qwen35 ,0 ,any ,RMS_NORM+MUL , 8
qwen35 ,0 ,any ,RMS_NORM+SCALE , 2
qwen35 ,0 ,any ,SSM_CONV+UNARY , 1
qwen35moe ,1 ,any ,ADD+ADD , 2
qwen35moe ,1 ,any ,GATED_DELTA_NET+CPY , 1
qwen35moe ,1 ,any ,MUL+ADD , 4
qwen35moe ,1 ,any ,RMS_NORM+MUL , 8
qwen35moe ,1 ,any ,RMS_NORM+SCALE , 2
qwen35moe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
qwen35moe ,1 ,any ,SSM_CONV+UNARY , 1
qwen3moe ,1 ,any ,MUL+ADD , 4
qwen3moe ,1 ,any ,RMS_NORM+MUL , 9
qwen3moe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
qwen3next ,0 ,any ,ADD+ADD , 2
qwen3next ,0 ,any ,GATED_DELTA_NET+CPY , 1
qwen3next ,0 ,any ,MUL+ADD , 4
qwen3next ,0 ,any ,RMS_NORM+MUL , 8
qwen3next ,0 ,any ,RMS_NORM+SCALE , 2
qwen3next ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
qwen3next ,0 ,any ,SSM_CONV+UNARY , 1
qwen3tts ,0 ,any ,RMS_NORM+MUL , 9
qwen3vl ,0 ,any ,RMS_NORM+MUL , 9
qwen3vlmoe ,1 ,any ,MUL+ADD , 4
qwen3vlmoe ,1 ,any ,RMS_NORM+MUL , 9
qwen3vlmoe ,1 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
qwen4exp ,0 ,any ,ADD+ADD+ADD , 1
qwen4exp ,0 ,any ,ADD+ADD+ADD+ADD+ADD+ADD+ADD , 9
qwen4exp ,0 ,any ,GATED_DELTA_NET+CPY , 1
qwen4exp ,0 ,any ,MUL+ADD , 4
qwen4exp ,0 ,any ,RMS_NORM+MUL , 13
qwen4exp ,0 ,any ,RMS_NORM+SCALE , 2
qwen4exp ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
qwen4exp ,0 ,any ,SSM_CONV+UNARY , 1
refact ,0 ,any ,RMS_NORM+MUL , 5
refact ,0 ,any ,RMS_NORM+MUL , 5
rnd1 ,0 ,any ,MUL+ADD , 4
rnd1 ,0 ,any ,RMS_NORM+MUL , 9
rnd1 ,0 ,any ,SOFT_MAX+ARGSORT+GET_ROWS+SUM_ROWS+CLAMP+DIV, 2
seed_oss ,0 ,any ,RMS_NORM+MUL , 5
smallthinker ,0 ,any ,MUL+ADD , 4
smallthinker ,0 ,any ,RMS_NORM+MUL , 5
smollm3 ,0 ,any ,RMS_NORM+MUL , 5
spark2_5 ,0 ,any ,RMS_NORM+MUL , 5
stablelm ,0 ,any ,NORM+MUL , 4
stablelm ,0 ,any ,NORM+MUL+ADD , 5
starcoder ,0 ,any ,NORM+MUL+ADD , 5
starcoder2 ,0 ,any ,NORM+MUL+ADD , 5
talkie ,0 ,any ,ADD+ADD , 2
xverse ,0 ,any ,RMS_NORM+MUL , 5