* CUDA: Handle compute type for NVFP4 on cublass path
Signed-off-by: ynankani <ynankani@nvidia.com>
* Use BF16 compute type for quantized models if HW allows
Signed-off-by: ynankani <ynankani@nvidia.com>
* Set acc prec to bf16 for nvfp4 as it needs atleast bf16 range
Signed-off-by: ynankani <ynankani@nvidia.com>
* Update ggml/src/ggml-cuda/ggml-cuda.cu
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* preserve op_params for per-expert matmul
Signed-off-by: ynankani <ynankani@nvidia.com>
---------
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.
We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)
* BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS
* AOCL-Blas : Add an AOCL-BLAS Quick Start and drop the fixed version path
* AOCL-BLAS doc : Note on ZenDNN
LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space
after every special token and parallel tool calls are separated by
<|end|>. The GPT-OSS handler rejects this output, so add a dedicated
handler, selected by the chat_format=llm-jp-harmony-v1 declaration in
the chat template.
Assisted-by: Claude Fable 5.1
* vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3
The original tokenizer configs for PLaMo-2 and PLaMo-3 have
`add_bos_token: true` and `add_eos_token: false` , but
_set_vocab_plamo() did not write the BOS/EOS metadata. The
PLAMO2 tokenizer path also ignored add_bos/add_eos during
tokenization.
Write the settings from tokenizer_config.json and honor them in
the PLAMO2 tokenization path. GGUFs without these keys keep the
previous behavior.
* Update conversion/base.py
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
The committed docs/ops/CPU.csv is out of sync with the current
test-backend-ops suite: 11 ops with CPU support (COL2IM_1D,
MUL_MAT_HADAMARD, SWIGLU_CLAMP, MUL_MAT_W4A4/W4A8, MUL_MAT_ID_W4A4/W4A8,
DSV4_HC_COMB/PRE/POST, LIGHTNING_INDEXER) are missing entirely, and
many other ops have fewer test cases than the suite generates now.
docs/ops.md (which CI requires to match the CSVs) therefore
understates CPU support.
Regenerated with:
test-backend-ops support -b CPU --output csv > docs/ops/CPU.csv
scripts/create_ops_docs.py
Note: ADD1 now reads unsupported on CPU because ggml_add1 is
GGML_DEPRECATED and the suite no longer generates test cases for it;
the CPU implementation itself is still present.
Assisted-by: Xing
#27941 disabled -sm tensor for qwen4exp because test-llama-archs asserted on the
Meta device once the fixture carried a PLE layer:
GGML_ASSERT(ggml_backend_buffer_is_meta(tensor->buffer)) at ggml-backend-meta.cpp:476.
With host-resident embeddings the PLE gather is a CPU node and hc_init (the REPEAT
that fans the embedding out to the hc streams) was first reached through layer 0's
PLE path, after that gather. ggml_backend_sched_split_graph pass 2 expands a device
assignment upwards only until it meets a CPU node, so the REPEAT stayed on the CPU
and the later reshape of hc_init inside the meta split viewed a host-resident node.
Expanding hc_init right after it is built puts the REPEAT directly before the first
device node, where pass 2 assigns it; the embedding reshape stays in the CPU split
and is copied in as a split input, as in deepseek4.
dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than that (an IQ4_NL row is only guaranteed to be a multiple of QK4_NL). Threads whose 32-value sub-block starts at or past k currently read and write past the end of the row. Skip those sub-blocks; for rows that are a multiple of QK_K the check never fires.
* hex-allreduce: add support for safe scatter mode
* hex-allreduce: pare down excessive comments
* hex-allreduce: re-write to remove register spills
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* llama: preserve original batch order for layer inputs
Assisted-by: Codex
* tests: cover layer-input order across KV layouts
Assisted-by: Codex
* tests: exercise layer-input ordering on CUDA devices
Assisted-by: Codex
* llama: make layer input reordering compatible with tensor split
Copy each microbatch tensor from offset zero and restore original row order after synchronization. Extend the layer-input regression to cover tensor split and repeated reads and decodes.
Assisted-by: Codex
* llama: restore token order for unmasked NextN embeddings
Use the original-token mapping for unmasked NextN rows, including when
layer-input capture is disabled. Keep masked NextN rows on the logits
output mapping and preserve offset-zero tensor copies.
Extend the existing regression to cover NextN alone, combined layer
capture, and masked outputs with repeated decodes and getters.
Validation: all 256 CPU/CUDA/tensor configurations pass. Qwen3.8-27B
Q4_K_M MTP completes MT-Bench at concurrency 16 before and after.
Assisted-by: Codex
* ggml: fix WebGPU reservation and OpenVINO hidden-state capture
Reserve WebGPU vector attention scratch across batch sizes and refresh reservations when NextN capture settings change. Preserve requested OpenVINO outputs, dynamic shapes, sequence counts, and current graph bindings.
Extend existing WebGPU regression coverage and enable strict allocation checks.
Assisted-by: Codex
* llama: defer regression test and backend fixes to follow-ups
Keep this PR focused on restoring token order for layer inputs and unmasked NextN embeddings. Remove the added regression test, OpenVINO and WebGPU changes, and the separate NextN reservation change.
Assisted-by: Codex
* llama: keep n_embd declaration in its original position
Assisted-by: Codex
* llama : pass token count to layer input extraction
Assisted-by: Codex
* llama : name original batch indices batch_idxs
Assisted-by: Codex
* llama : name extracted embedding indices embd_batch_idxs
Assisted-by: Codex
* llama : tag target embedding reordering
Assisted-by: Codex
* llama : tag extraction and name the index capture flag
Assisted-by: Codex
After the device decode, flip causal_attn off, decode n_ubatch/2 then
n_ubatch tokens. Both have the same node count, so a shape that depends
on the flag makes the second reallocate at an unchanged graph size,
which aborts under GGML_SCHED_NO_REALLOC. Skipped for the encode archs.
* convert: fix LoRA conversion crash for Qwen3.5 V-head reorder
_reorder_v_heads does reshape+permute+reshape to reorder V heads from
grouped to tiled order. LoraTorchTensor.reshape() cannot split its
row dimension (A matrix), so converting Qwen3.5 LoRA adapters that
target out_proj crashes with NotImplementedError.
Fix: detect LoRA tensors and apply the equivalent index permutation
directly — column reorder (dim=last) permutes A's columns, row
reorder (dim=0) permutes B's rows. This is mathematically identical:
(B @ A)[:, perm] == B @ A[:, perm]
(B @ A)[perm, :] == B[perm, :] @ A
Verified: both paths produce exactly zero diff against the full-tensor
reorder on random (rank=32, 4096×4096) matrices.
Fixes#21125
Signed-off-by: Radu Swigler <radu@swigler.com>
* convert: add ty: ignore for hasattr-guarded LoRA call
Assisted-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix comment
* nowrap
---------
Signed-off-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Radu Swigler <radu@swigler.com>
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Assisted-by: Claude Opus 4.6 <noreply@anthropic.com>
The sparse indexer mask is built with a set_rows scatter. Padded pools,
absent sequences and missing tail cells all pointed to the same n_kv
sentinel row, and invisible pools picked by top_k to fill the selection
overlap the tail cells of the token, so several CPU threads wrote the
same element (ThreadSanitizer data race in the sanitize CI).
Allocate the slot mask for both selection paths and route every dead
slot to its own dump row n_kv + slot. Live slots address disjoint cells,
so the scatter indices of a token are unique.
* ggml: fix integer overflow guard for zero-element tensors
* ggml: validate number of elements in tensor to prevent integer overflow
* ggml: fix error print
* cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast
On Windows the simple input reader sends CTRL_C_EVENT to every process
attached to the console when stdin reaches EOF, killing unrelated
processes such as a supervising agent. The CLI only stopped on EOF
because of that self inflicted SIGINT; on POSIX, and with the advanced
reader, it spins forever printing prompts.
Drop the broadcast so both platforms just return an empty read, and
treat an empty read as EOF in the chat loop and the model selection,
since a submitted line always ends with a newline.
* cli: keep the newline of a trailing "/" and stop mtmd-cli on EOF
A lone "/" came back as an empty read and was taken for EOF, and
mtmd-cli only stopped on EOF through the removed broadcast.
The fixture recycles its two blocks over 8 cache slots, so the fp16
error builds up past the 1e-4 NMSE bound on the Vulkan T4 and WebGPU
jobs of Models Backend. Two l-cycles keep every branch of the cycle
loop and halve the error.
* llama: llama_prefetch_rows
* llama: support row prefetch on Windows
Apply the Windows port contributed by @praneshgo unchanged.
Source: https://github.com/ggml-org/llama.cpp/pull/29599#issuecomment-5887721014
* avoid exposing llama-mmap in model code, route via llama-impl
* add windows check, only prefetch in lazy mode
* cont : clean-up
* cont : fix build
* cont : clarify padding token for gemma4
---------
Co-authored-by: Pranesh Gonegandla <pranesh.iitp@gmail.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* ggml : add BF16 unary, GLU, binary and scale ops (CPU, CUDA)
* ggml-cpu : use per-op _bf16 functions for BF16 unary and GLU ops
Assisted-by: Claude Opus 5.5
* CUDA: use ggml_cuda_cast in binbcast and unary kernels to fix the HIP bf16 build
* ggml-openvino : reject BF16 SCALE and mixed-type BF16 ADD/MUL/SUB
* cpu: accept BF16 in src1 of mul_mat
ggml_conv_1d_dw builds its im2col in F32 when the kernel is BF16, then
calls ggml_mul_mat(im2col, kernel), which puts F32 in src0 and BF16 in
src1. The CPU backend refused that combination, so it was reported as
unsupported on every backend and never compared against anything.
Widen BF16 into the F32 work buffer, next to the existing packing of F32
into vec_dot_type. This is the arithmetic the Metal mat vec kernel
already uses, both operands promoted to float and accumulated in float,
so the two agree exactly rather than approximately.
Cover it with a conv_1d_dw test over F32, F16 and BF16 kernels, plus
three mul_mat cases with BF16 in src1.
* vulkan: reject BF16 in src1 of mul_mat unless src0 is BF16
supports_op only checked the src1 type for non contiguous tensors, so
a contiguous BF16 src1 was accepted and the pipeline lookup asserted.
The only BF16 src1 path is the BF16 x BF16 multiply, every other src0
type now reports the op as unsupported and the scheduler keeps it on
the CPU.
The BF16 kernel case of the conv_1d_dw test needs the f32 x bf16
mat vec variants of the Metal backend, which land separately.
This commit adds an optional --add-bos token command line option to the
run-org-model.py script.
The motivation for this is that there are models, for example Gemma4,
that explicitely set the add_bos value to true in llama-vocab.cpp even
if the original model does not set this value to True.
It would be nice to be able to force the models to agree on the bos
token so that logit verification can proceed.
Refs: https://github.com/ggml-org/llama.cpp/pull/21500
* ui : shared model display primitives
Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.
Assisted-by: pi:GLM-5.3-Flash
* ui : remember hub avatars that failed to load
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : render shared model row hints as native titles
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : fix badge guard for draft sidecars, keep parameter precision
hasBadges now counts draft sidecar badges, so a sidecar-only model still
renders. Billions keep one decimal for hub counts and stay bare for whole
values. Avatar failures track the org instead of the instance, and the
download progress bar no longer pulses while determinate.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model download pipeline
Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.
Assisted-by: pi:GLM-5.3-Flash
* ui : route sidecar tag lookup through the sidecars util, validate the paused list
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model memory-fit estimation
Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.
Assisted-by: pi:GLM-5.3-Flash
* ui : cover the memory-fit and tool-use heuristics in tests
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : Hugging Face Hub data layer
Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.
Assisted-by: pi:GLM-5.3-Flash
* ui : strip provider tilde prefix from hub avatar urls
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : trim redundant comments in the HF data layer service
Per review: drop JSDoc that restates the method name and inline comments
that restate the code; keep only comments carrying non-obvious context.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : harden the HF data layer error typing, cover the helpers in tests
Carries the HTTP status on retryable fetch errors instead of matching the
message text. Marks expand-dependent catalog fields optional and documents
the data/models index pairing. Adds table tests for the pure helpers.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model id grammar for sidecars, quants and capability parsing
Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.
Assisted-by: pi:GLM-5.3-Flash
* ui : escape sidecar tokens in the regex alternation
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : type-safe API types, fetch helpers and download-ready models store plumbing
Assisted-by: pi:GLM-5.3-Flash
* ui : document the model list index pairing, fix an em-dash
Assisted-by: pi:zai-org/GLM-5.3-Flash
* Update tools/ui/src/lib/components/app/chat/index.ts
Co-authored-by: Pascal <admin@serveurperso.com>
---------
Co-authored-by: Pascal <admin@serveurperso.com>
* openvino: serve GET_ROWS on a weight view from the base Constant
Resolve view_src when collecting weight Constants so a view over a
quantized weight no longer becomes a dynamic typed Parameter, and fold
the row offset of the view into the gather indices instead of slicing
the dequantization subgraph.
* openvino: lift the quantized GET_ROWS view rejection
The supports_op rejection of a quantized src0 view with a nonzero
offset keeps the vs0 GET_ROWS cases of #28253 away from OpenVINO.
The weight view now resolves to the base Constant with the row offset
folded into the gather indices, so the rejection goes away.
#27773 adds the glm5-next arch without its rows in the Metal fusion
baseline, so test-fusion --check fails on it. The rows come from
test-fusion --record on an M5 Max, and --check passes 270/270.