* Make the drafter probabilistic and the target verify by rejection sampling
* Drop stale spec_draft_q before drafting
* Fallback to argmax sampling for grammar-constrained requests and adding flag for enabling probabilistic draft sampling. Default flag value is greedy.
* Support grammar-constrained requests in rejection sampling
* Fix - renormalize distribution after masking
* copy rng on sampler copy and re-accept drafted tokens on replay
* Fix draft sampler sharing the target's rng stream
* Simplify the rejection sampler's inputs and move replay to the server
* Truncate the draft candidates along with the draft
---------
Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com>
Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>
* ggml-quants : avoid invalid rounding in qkx3 scale search
The imatrix scale search can produce an infinite, NaN, or otherwise out-of-range value when the fitted minimum collapses to the maximum or makes the range extremely small. That value is then passed to nearest_int and can trip its assertion in Debug builds.
Clamp the quantization level to [0, nmax] before rounding so valid in-range values behave the same as before while invalid scale-search results no longer reach nearest_int.
Add regression coverage for degenerate imatrix groups across q2_K, q4_K, q5_K, q4_1, and q5_1.
Fixes#29804.
Assisted-by: Claude Opus 5.5
* tests: print degenerate imatrix quant types
* ggml-cpu : fix soft_max_back wrong output when dst aliases src1
GGML_OP_SOFT_MAX_BACK is listed in ggml_op_can_inplace, so the graph
allocator may assign dst to alias either src0 (dy) or src1 (y).
The result was built in several steps:
ggml_vec_cpy_f32 (nc, dx, dy);
ggml_vec_acc1_f32 (nc, dx, -dot_y_dy);
ggml_vec_mul_f32 (nc, dx, dx, y);
ggml_vec_scale_f32(nc, dx, scale);
When dst aliases src1, the first step overwrites y and the third step
then reads the overwritten values, so the output is silently wrong.
Aliasing dst with src0 is unaffected. The CUDA kernel completes its
reduction before writing and is already safe.
Replace the sequence with a single fused loop that reads both sources
before writing, which is correct under either aliasing.
Add a regression test that marks dy as a graph output so the allocator
is forced to alias dst with y, asserts that the alias actually
happened, and compares against values computed on the host.
* cont : remove comment
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* metal : add tensor API flash attention kernel for F16 KV
* cont : add tensor FA kernels for DK=DV=512 and DK=576, DV=512
* cont : support attention sinks, ALiBi and logit softcap in the tensor FA kernel
* cont : add tensor FA kernel for DK=192, DV=128
* init conversion
* convert: ok
* model loaded
* add server code
* improve conversion script
* support shared prompt prefix
* add docs, imorove UX a bit
* add vision support
* add openjev tiny model for testing
* add dev docs
* support lev & kev
* clean up
* fix lev noul
* fix py lint
* nits docs
* clarify about not supporting date_facts
* Adding wide-load mmvq for Q8_0 and esimd dmmv for q8_0
Assisted-by: Codex
* remove guard for q8_0
* remove docs
* Simplify by committing to clean code without fallback
* Add feature flag as requested
Assisted-by: Claude Opus 5
---------
Co-authored-by: cwriter <cwriter@localhost>
* ggml : add `alloc_buffer_n` to buffer type interface
Add alloc_buffer_n method to ggml_backend_buffer_type_i
interface, with a public API ggml_backend_buft_alloc_buffer_n.
- Default implementation in ggml-backend.cpp handles multi-buffer
splitting and tensor allocation via ggml_tallocr
- Meta buffer type provides custom implementation that creates
per-device sub-contexts and delegates to simple buffer types
- ggml_backend_alloc_ctx_tensors_from_buft now collects tensors
into a list and delegates to the new API
- Remove temporary ggml_backend_meta_alloc_ctx_tensors_from_buft
- Add NULL alloc_buffer_n to all existing buffer type
interfaces (cpu, metal, openvino, hexagon, webgpu, zdnn, virtgpu, repack)
Assisted-by: llama.cpp:local pi
* cont : fix `cur_buf_size` init after flushing a buffer
* ggml : add TODO tag for shared buffer split logic
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
* tests : add alloc_buffer_n coverage
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
* cont : fix compile warnings
* tests : add descriptions for alloc_buffer_n tests
Assisted-by: pi:llama.cpp/Qwen3.8-27B
* ggml : address review comments on alloc_buffer_n
- restore GGML_LOG_ERROR on buffer alloc / tensor init failure in the
default impl (name the failing tensor)
- check the malloc result and drop the _impl indirection in
ggml_backend_alloc_ctx_tensors_from_buft
- remove comments that restate the code
- fix the TAG_ALLOC_SHARED_BUFFER_SPLIT typo
Assisted-by: pi:llama.cpp/Qwen3.8-27B
* ggml : add get_alloc_size_n to buffer type interface
- Add ggml_backend_buft_get_alloc_size_n public API
- Add optional get_alloc_size_n callback to ggml_backend_buffer_type_i
- Share tensor->buffer planning between alloc_buffer_n default and get_alloc_size_n default
- Replace unchecked realloc with std::vector in alloc_buffer_n default
- Make ggml_backend_alloc_ctx_tensors_from_buft_size use the new API
- Add test-alloc coverage for get_alloc_size_n
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
* cont : report malloc failure
In tool.uv.sources, torch was unconditionally pinned to the custom
pytorch CPU index, which lacks macOS Darwin wheels and causes uv sync
to fail on macOS. Add the sys_platform == 'linux' marker to match the
existing Poetry dependencies configuration.
Assisted-by: Antigravity
Resolves: https://github.com/ggml-org/llama.cpp/issues/29176
* hexagon: add q2_k and q3_k quant type support
* hex-qk: consistent allocation of src1_row_size
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* tests : simplify function signature
* llama : clamp kpool re-pool bound to existing pools
The n_tokens/kpool + n_seqs_unq bound on n_new_g overshoots when a batch
fills the whole cache: n_ctx tokens complete exactly n_ctx/kpool pools, so
the +1 pads new_pool_idxs/new_pool_rep one entry past n_pool_real. Graph
reserve only covers n_pool_real entries, so the first full-context decode
builds bigger tensors than reserved and ggml-alloc demands a graph
reallocation (abort under GGML_SCHED_DEBUG_REALLOC=1).
Clamp the bound to n_pool_real: a ubatch can never mark more pools than
the cache holds, and reserve's n_pool_max already covers that.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
* cont : cap to n_pool_max
* hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA
* hex-cpy: various fixes on top of the concat optimizations
Removed CONCAT_DMA_MIN_ROW logic, it was broken with 64-bit DMA.
While it's kinda silly to use DMA for tiny stuff if that tensor gets mapped to an extended buffer the only way to read it is DMA.
Added missing dma_queue_flush() calls.
Added additional guards for conditions we don't support.
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* convert : write Gemma embedding scale for DFlash drafts
A DFlash draft shares the target's token embeddings. Gemma scales them by sqrt(hidden_size) in the forward pass, and the draft config does not state that scale, so the converted draft read unscaled embeddings.
Take the scale from the target config when the draft config has none.
Assisted-by: Claude
* convert : check with get_model_architecture for gemma models
* cuda : route sm70 to the Turing MMVQ nwarps table
Volta (sm_70) has no MMVQ parameter table of its own and falls through
to GENERIC, which launches K-quant batch-1 decode (ncols_dst == 1) at
nwarps=4. sm_70 shares TURING's tuning: the K-quant vec_dot prefers
nwarps=2 there. Route sm_70 to the existing MMVQ_PARAMETERS_TURING
table in both the device and the host table selector.
Measured on one Tesla V100 32GB PCIe (PG500-216, driver 580.178.04,
CUDA 12.0.140) with Qwen3.8-27B Q4_K_M, tg128, interleaved A/B in 6
ABBA blocks with paired per-block deltas: +1.091 t/s = +3.17 %
(t = +49.0, all six per-block deltas positive); perplexity
bit-identical (6.3697 +/- 0.04066 both builds, wiki.test.raw). The
patched build's K-quant mul_mat_vec_q kernels launch at nwarps=2
(cubin EIATTR_MAX_THREADS) while Q4_0/Q8_0 stay at nwarps=4, and the
same measurement on the September master base gave +3.84 % (t = 85).
The tuning originates from the V100-focused fork anyei/llamacpp-v100
(MIT), commit b912d1b1e, which carries a dedicated
MMVQ_PARAMETERS_VOLTA table; a cubin-level comparison confirmed that
routing sm_70 to the existing TURING table is equivalent for the
K-quant batch-1 path this change affects, so this is the minimal
2-line form. https://github.com/anyei/llamacpp-v100/commit/b912d1b1e
Original-patch-by: anyei <angelyoelroblesmercedes@gmail.com>
* Update ggml/src/ggml-cuda/mmvq.cu
---------
Co-authored-by: tkittich <tkittich@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* metal : release temporary private transfer buffers
Assisted-by: OpenAI Codex
* metal : fix order and formatting
---------
Co-authored-by: Niklas Wenzel <dev@nikwen.de>
* CUDA: Handle compute type for NVFP4 on cublass path
Signed-off-by: ynankani <ynankani@nvidia.com>
* Use BF16 compute type for quantized models if HW allows
Signed-off-by: ynankani <ynankani@nvidia.com>
* Set acc prec to bf16 for nvfp4 as it needs atleast bf16 range
Signed-off-by: ynankani <ynankani@nvidia.com>
* Update ggml/src/ggml-cuda/ggml-cuda.cu
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* preserve op_params for per-expert matmul
Signed-off-by: ynankani <ynankani@nvidia.com>
---------
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.
We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)
* BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS
* AOCL-Blas : Add an AOCL-BLAS Quick Start and drop the fixed version path
* AOCL-BLAS doc : Note on ZenDNN
LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space
after every special token and parallel tool calls are separated by
<|end|>. The GPT-OSS handler rejects this output, so add a dedicated
handler, selected by the chat_format=llm-jp-harmony-v1 declaration in
the chat template.
Assisted-by: Claude Fable 5.1