Commit Graph
18 Commits
Author SHA1 Message Date
649dcb1036 add GLM-5.3-Flash (GLM5-Next) support (#27773)
* Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx

* Add initial MTP support

* Merge branch optimizations. Reduce allocated compute buffer size, speed up long context decode, fla, and slight MTP improvements.

* Review driven changes, remove env vars, protect tensors

* Strip MTP for initial PR

* Clean up after mtp strip

* Clean up after mtp strip

* Update speculative.cpp

* Update llama-context.h

* Clean up after mtp strip

* Fix tokenizer ignore merges

* Improve quantization protection selection

* Refactor mhc helpers, graph base

* Lint Fixes

* Apply suggestions from code review

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

* Skip glm5-next in model saver, fix CRLF

* Skip glm5-next in sweep

* Remove T4 fallback

* Review cleanup

* Review suggestions

* Defer separate MTP gguf handling to MTP PR, drop filter

* Repad n_head_kv

* kpool init apply

* Order by descending score

* Drop guard

* read kpool from hparams, clarify kpool cache flags, remove kpool_build_state(nullptr)

* Add glm5-next support to model saver and add arch test fixture

* Review cleanup

* Kpool pooled caching clarify

* Add multi stream support

* Finish Rebase

* Sparse FA fir DSA prefill

* Const

* Update llama-model.cpp to fix rebase error

* gguf-py : merge tensor map entries for HC tensors

* model : use build_gdn_l2_norm in GLM5_NEXT implementation

* chore : remove trailing whitespace

* model : use new OP precision setting API in GLM5_NEXT implementation

* mtmd : use ggml_swiglu_clamp in GLM5V and apply the image token limit

The two clamps around swiglu_split are what ggml_swiglu_clamp already does,
so the clamp bounds collapse back to one value. GLM5V also never called
set_limit_image_tokens(), so --image-max-tokens had no effect.

Assisted-by: Claude Opus 5
(cherry picked from commit 46d18e12d422be4cc04a70e4a9a9e0168bb3d5b7)

* llama : keep the GLM5-Next k-pool layout across ubatches

The layout was rebuilt from a full cell scan on every ubatch. Pools are fixed
by the positions relative to the sequence's first one, so the layout now lives
on the memory and a ubatch only appends to it.

A sequence edit no longer stales every pooled key either, only the ones at or
after the edited position, which makes a tail seq_rm free. The pooling subgraph
is built unconditionally so the graph shape no longer changes every kpool
tokens, and the pool axis is folded into rows before soft_max, which otherwise
exceeds the CUDA gridDim.y limit past n_kv 262144.

Assisted-by: Claude Opus 5
(cherry picked from commit 5d1c40b93e17fddbf73b785efe43e0d02ccb3977)

* model : write the GLM5-Next recurrent rollback checkpoints

The conv state and the delta net state were only written to the live row, so a
rollback restored whatever the checkpoint rows happened to hold. Take the same
route as kimi-k3: build_recurrent_attn for the state, and write all K_rs conv
groups. That also drops a state view that assumed contiguous rows.

Enroll the arch in test-recurrent-state-rollback, which catches this under its
garbage-filled cache pass.

Assisted-by: Claude Opus 5
(cherry picked from commit 5ace37e86d5d448e83ef5dde5632c748185b18cd)

* llama: fix PR #27773 test-save-load-state restore failure

Clear the attention and indexer cache data after a failed hybrid state restore so restored NaNs cannot affect a later sequence.

Assisted-by: Codex

* llama: fix PR #27773 gpu-rocm graph reallocation

Reserve the full GLM5-Next pool capacity and dirty pool count. The gpu-rocm Test step aborts when n_new grows while the graph node count stays fixed; CUDA, Vulkan, Metal, and WebGPU checks report the same error.

Assisted-by: Codex

* llama : fix GLM5-Next k-pool layout staleness after edits and shared teardown

Two defects in the cross-ubatch k-pool layout added by the k-pool commit:

1. Wrong results. An edited sequence only rebuilt its pool layout when its cell
   count changed, so if the first ubatch after an edit added back exactly as many
   cells as were removed, the stale position-to-cell list survived. With a unified
   cache and more than one sequence, where another sequence takes the freed cells,
   the reused layout points at the wrong cells (CPU: large logit drift, CUDA: NaN).
   Rebuild whenever the sequence is stale, not only on a size mismatch.

2. Slowdown. "shared" mode was assumed to end only with an edit that forces a
   rebuild, but sharing also ends when the other sequence is removed. The survivor
   kept shared = true, pinning cache_safe off and re-pooling every pool on every
   ubatch (server trigger: n>1 completions with -kvu, via the seq_cp in
   copy_state_to). In seq_rm, if the layout has shared cells, stale every sequence
   so one rebuild re-derives sharing and cache_safe returns to 1.

Assisted-by: Claude Opus 5

* llama : fix build_attn_mha stream stride for non-contiguous q

build_attn_mha split the batch into streams with a stream stride of
q->nb[3]/n_stream. That only equals one stream's span, (ne[2]/n_stream)*nb[2],
when q is contiguous. GLM5-Next is nope-only, so it does not concat a rope part
and passes the permuted q_absorbed straight in, where nb[3] != ne[2]*nb[2]; the
stride was then n_head times too large and every stream s >= 1 read another
head's queries. Split-KV (-np N without --kv-unified) multi-stream prefill was
wrong for every stream past the first. Unified KV and decode were unaffected
(n_stream == 1, and decode takes the gather path). Other MLA models concat rope
so q is contiguous and the computed value is unchanged for them.

Compute the stride from the token dimension, which is identical for a
contiguous q.

Assisted-by: Claude Opus 5

* llama : re-derive GLM5-Next k-pool sharing on state_read/state_drop

The shared-cell teardown added to seq_rm (stale every sequence when the layout
has shared cells, so a survivor does not keep shared = true and pin cache_safe
off) was missing from the other paths that can free shared cells: state_read
and state_drop staled only the one sequence. Apply the same re-derivation there
and correct the comment that claimed sharing ends only via an edit or seq_rm.

Assisted-by: Claude Opus 5

* quant : drop duplicate GLM5-Next hc_ filter

The hc_ name filter was listed twice in the GLM5_NEXT protection block.

Assisted-by: Claude Opus 5

* glm5-next: scope K-pool cache access to indexed operations

* glm5-next: keep K-pool access in hybrid index memory

* glm5-next: keep mHC graph builders model-local

* glm5-next: mark only touched pools per ubatch

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Piotr Wilkin <ilintar@gmail.com>
2026-09-30 14:20:32 +08:00
Xuan-Son Nguyen 163a40796f model, mtmd: fix gemma4 vision handling (#28335)
* model, mtmd: fix gemma4 vision handling

* nits
2026-09-04 12:23:27 +02:00
Yaniss Amazouz c61b98b875 model: add NVIDIA Nemotron-3-Puzzle-75B-A9B (NemotronHPuzzle) support (#25444)
* hparams: add per-layer n_ff_exp/n_expert_used arrays with scalar-or-array loading

G1/G2 infrastructure for variable-per-layer expert FFN size and top-k routing
(required for Puzzle-75B which has 5 distinct n_ff_exp values and 7 top-k values
across its 40 MoE layers).

Design: rename scalar members to _impl suffix (following existing convention),
add LLAMA_MAX_LAYERS arrays, add n_ff_exp(il)/n_expert_used(il) accessors with
scalar fallback. No new GGUF keys: reuses existing expert_feed_forward_length and
expert_used_count keys via get_key_or_arr (scalar -> broadcast, array -> per-layer).

- llama-hparams.h: n_ff_exp -> n_ff_exp_impl, n_expert_used -> n_expert_used_impl;
  add n_ff_exp_arr / n_expert_used_arr arrays; add per-layer accessor declarations.
- llama-hparams.cpp: implement n_ff_exp(il) and n_expert_used(il); out-of-range
  il returns impl safely (shared code, no abort).
- llama-model.cpp: central n_expert_used load changed to get_key_or_arr; derive
  impl as max-of-array for validations and backward compat; zero both new arrays;
  HunyuanVL override also zeroes n_expert_used_arr.
- llama-graph.cpp: aggregation loop in build_moe_ffn uses hparams.n_expert_used(il)
  so per-layer top-k bounds the ggml_view loop correctly.
- All other files: mechanical rename hparams.n_{ff_exp,expert_used} -> *_impl.
  Scalar arches are unaffected (broadcast fills all array slots with the single value).

(cherry picked from commit 269a81e03d)

* nemotron-h: use per-layer n_ff_exp(il) and n_expert_used(il) at MoE call-sites

Load n_ff_exp via get_key_or_arr into hparams.n_ff_exp_arr in load_arch_hparams;
derive impl as max for existing uniform GGUFs.

In load_arch_tensors, compute n_ff_exp_i = hparams.n_ff_exp(i) with fallback to
n_ff(i)/n_expert_used(i) for GGUFs that omit expert_feed_forward_length.

In build_ffn_layer, pass hparams.n_expert_used(il) to build_moe_ffn so per-layer
top-k is used for expert routing selection.

All other nemotron-h behaviour (mamba2, attention, shared-exp, latent projection,
routed_scaling_factor, expert_weights_norm, sigmoid gating) is unchanged.

(cherry picked from commit b1878a1017)

* arch/*.cpp + gguf-py: mechanical rename n_ff_exp->n_ff_exp_impl, n_expert_used->n_expert_used_impl

All non-nemotron arch files continue using the scalar impl member directly.
Behaviour is identical: the impl value is the broadcast value from the GGUF scalar.

gguf_writer: add_expert_feed_forward_length and add_expert_used_count now accept
int | Sequence[int], mirroring add_feed_forward_length, so converters can write
per-layer arrays with the same existing GGUF keys.

(cherry picked from commit 8f009f54be)

* convert: support NemotronHPuzzleForCausalLM (per-block MoE config)

Parse block_configs/mtp_block_configs into per-layer arrays (scalar-or-array
keys), append the MTP [attention, moe] sub-blocks as blk.88/blk.89 with
nextn tensors, accept the backbone.* prefix, and register the arch.
Also fix a pre-existing undeclared _experts attribute on NemotronHModel.

(cherry picked from commit d1a592f278)

* nemotron-h: distinguish Nemotron 3 Puzzle (75B.A9B) from Super (120B.A12B)

Both have 88 layers; the per-layer expert_used_count array (heterogeneous
for Puzzle, broadcast-uniform for Super) is the discriminator.

(cherry picked from commit f824e09dc8)

* convert: accept the official Puzzle BF16 checkpoint's tensor naming

The officially distributed BF16 checkpoint (NVIDIA-Nemotron-Labs-3-Puzzle-
75B-A9B-BF16) names the trunk model.* (model.layers.*, model.embeddings,
model.norm_f) where the original release used the NemotronH-style
backbone.*, and spells the router bias e_score_correction_bias instead of
e_score_correction.bias. Normalize both at the top of
NemotronHPuzzleModel.modify_tensors so either checkpoint converts; every
tensor name in the official index (42683 keys, MTP head included) resolves
through the tensor map after normalization.

(cherry picked from commit 189b67fc2c)

* laguna: use n_ff_exp_impl for the uniform-MoE FFN size

Laguna landed after this branch was cut and reads hparams.n_ff_exp as a
scalar. This series turns it into a per-layer array with an n_ff_exp(il)
accessor, so the three scalar reads no longer compile. Laguna is a
uniform MoE, so point them at the scalar fallback n_ff_exp_impl, same as
deepseek2/qwen3moe/gemma4 in this series. No behaviour change.

(cherry picked from commit dbedc9e19c)

* arch: extend the n_ff_exp/n_expert_used rename to archs added upstream

kimi-k3, dflash, bailingmoe3, deepseek4, granite-swa and the nemotron-h MTP
block still referenced the scalar fields by their old names. n_ff_exp and
n_expert_used are accessors now, so those reads no longer compile; point the
non-per-layer archs at the _impl scalars and use the indexed form where the
call site is per-layer.

* convert: keep Puzzle opted out of the NemotronH MTP export path

#26725 added MTP export to NemotronHModel, keyed on num_nextn_predict_layers.
Puzzle's config carries that key, but NemotronHPuzzleModel bypasses
NemotronHModel.__init__ (its per-block config needs a different setup), so
_mtp_bid was never assigned and modify_tensors raised AttributeError on any
mtp.* tensor. Puzzle's head is also laid out by mtp_block_configs, not the
mtp.layers.* form the base maps.

Set _mtp_bid to None, drop mtp.* in filter_tensors, and declare
supports_mtp_export = False so --mtp / --no-mtp fail at the CLI.

* llama: replace n_ff_exp/n_expert_used scalars with per-layer accessors

Follow-up to review feedback: the previous revision kept the scalar
hparams fields alongside the new per-layer arrays, which duplicated
state that get_key_or_arr already handles by broadcasting a scalar
value over every layer.

Drop both scalars and expose n_ff_exp(il) / n_expert_used(il) built
exactly like the existing n_head_kv(il) and n_ff(il) accessors: they
index the array and GGML_ABORT out of range, with il defaulting to 0
so genuinely uniform call sites stay a plain n_ff_exp().

Arch loaders now read both keys through get_key_or_arr over
n_layer_all, and the n_expert_used validation checks the maximum
across layers instead of a single field.

* llama: restore per-key required flags on the expert hparam reads

The scalar-to-array conversion passed required=false at every call site,
which silently made mandatory keys optional. Each read now carries the
same required flag it had before the conversion.
2026-09-03 08:53:08 +02:00
Xuan-Son Nguyen 9400c8946e model: correctly support input vision for deepseek4 (#28154)
* model: correctly support input vision for deepseek4

* nits
2026-09-02 19:14:46 +02:00
Aman Gupta 8e93a9773b CUDA + ggml: add sparse-fa for DSV4/GLM (#27970) 2026-09-02 17:27:37 +03:00
Sigbjørn Skjæret 9d817213a0 model : load hparams.n_layer_nextn before n_layer() calls (#28159)
* load hparams.n_layer_nextn before n_layer() calls

* remove duplicate loads
2026-09-01 13:55:45 +02:00
Aman GuptaandGeorgi Gerganov b0539c43ed DeepseekV4: fix rollback with multi-seq (#26756)
* DeepseekV4: fix rollback with multi-seq

* fix model loading

* make pending rollback single use

* only clear cache for seq_id for full load

* add assert for compress ratio

* make graph topology static

* pass true instead of flags in clear_compressed

* cont : clean-up + TODOs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-08-23 13:57:49 +03:00
Xuan-Son Nguyen 873e5d8e39 model: use ggml_rope_set_offset() (#27382)
* model: use ggml_rope_set_offset()

* partially apply to deepseek2
2026-08-21 18:54:29 +02:00
Georgi Gerganov 929d47a391 graph : create V as a view of K in the k_iswa build_attn (#27392)
build_attn with the llm_graph_input_attn_k_iswa input was using the cached K
tensor itself as V. Create V as a view of K (the first v_cur->ne[0] elements
of each row), like the other K-only build_attn overloads.

The deepseek4 MTP call site now passes the kv tensor as v_cur.

Assisted-by: pi:llama.cpp/Qwen3.8-27B
2026-08-20 10:00:35 +03:00
Georgi Gerganov 1269cb1ff1 model : allow reshape of tensors during load (#26531) 2026-08-04 09:06:44 +03:00
Aman Gupta 596a5795bd DeepseekV4 MTP + DSpark (#25784) 2026-08-02 20:55:34 +08:00
Georgi Gerganov dee2a846b8 ggml : adjust logic for offloading ops to weight's backend (#25832)
* ggml : adjust logic for offloading ops to weight's backend

* llama : dsv4 graph fixes
2026-07-27 14:54:46 +03:00
Aman Gupta 0dc74e332e DeepseekV4: Add fused hyper-connection ops (#25585)
* dsv4 hc-ops

* add missing files;

* add cparams

* update rpc version

* address review comments

* address review comments
2026-07-17 00:33:33 +08:00
Aman Gupta 33a75f41c3 DeepseekV4: reduce graph splits (#25702) 2026-07-15 15:47:18 +08:00
fairydreamingandStanisław Szymczyk 00f5442cc4 ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer (#24231)
* ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer

* ggml : remove scale parameters from lightning indexer OP, add f16 mask parameter

* tests : add GGML_OP_LIGHTNING_INDEXER tests

* ggml : bump RPC version

* chore : check if lightning indexer input tensors are not transposed

* tests : count flops instead of bandwidth in lightning indexer test

* chore : add missing const

* chore : whitespace

* ggml : renamed variables in CPU lightning indexer implementation

* ggml : fix lightning indexer mask broadcasting

* tests : tests for lightning indexer mask broadcasting

* chore : whitespace

* llama : use GGML_OP_LIGHTNING_INDEXER in DeepSeek V3.2 and DeepSeek V4 models

---------

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
2026-07-11 11:39:07 +02:00
fairydreamingandStanisław Szymczyk 2ed3c1abbb llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 (#25370)
* llama : make all KQ masks (except the lightning indexer one) f16 if FA is used and remove zero attention bias in DeepSeek V4

* llama : remove dead code that repeats unified raw_k cache for each stream in DeepSeek V4 - no longer needed as raw_k is always non-unified.

---------

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
2026-07-10 09:06:58 +02:00
Aman Gupta 024c46ae4e llama: fix quantized kv-cache for dsv4 (#25202) 2026-07-07 17:46:57 +08:00
8c146a8366 DeepSeek V4 (#24162)
* convert: add dsv4 conversion

* add basic setup

* add llm_graph_input_dsv4

* add save-load state

* add sinkhorn eps - correction by @fairydreaming

* add rope fix

* cleanup dead code

* fix bugs

* support pro model: added by @fairydreaming

* remove redundant V cache

* Chat template

* remove debugging leftovers

* Add mechanism for inlining templates based on architecture

* s/deepseek-v4-flash/deepseek4/g

* s/deepseek-v4-flash/deepseek4/g continued

* enable graph reuse

* enable FA

* fix test llama archs

* rename

* compatibility with antirez ds4 GGUFs

* simplified set_gguf_parameters() by calling super class method, replaced moe.score_func with expert_gating_func.

* reserve worst-case kv-cache

* revert max split inputs

* address review comments

* add padding to enable FA

* pad only the final value of plan.n_kv to 256

* remove built-in cpp chat template

* cont: remove cpp built-in template

* rm outdated test

* replace ggml_view_3d() with ggml_reshape_3d()

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* only support n_seq=1 for now

* remove unused var

* cont: remove unused var

* use scale bias

* use correct ptr for can_reuse

* remove gen-chat-inline-templates.py

* simplify graph reuse

* cont: cleanup

* remove unused inputs

* enable partial checkpointing

* add correct shape for kq_mask + set llama_model_n_swa to 0 for dsv4

* precompute source_idx + add comment about dummy write

* support multi-seq

* remove restored_trim_pos

* use split_equal when possible

* fix indent

* address review comments

* use LLM_KV

* fix ci

---------

Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: fairydreaming <166155368+fairydreaming@users.noreply.github.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-06-29 16:58:51 +08:00