Commit Graph
10257 Commits
Author SHA1 Message Date
Pedro Cuenca 4037e13ea9 Merge branch 'onyx' into onyx-vision 2026-08-08 09:40:03 +02:00
Pedro Cuenca fe54c4f6f6 Merge pull request #9 from huggingface/onyx-swa-rope-theta
onyx: use the model rope theta on sliding-window layers
2026-08-07 19:19:29 +02:00
Daniel Han 46f2dc52df onyx: use the model rope theta on sliding-window layers 2026-08-07 15:35:34 +00:00
Pedro Cuenca 1f2a51f268 build_vit 2026-08-06 23:49:35 +02:00
Pedro Cuenca ead38d271f Remove duplicated function 2026-08-06 20:40:05 +02:00
Pedro Cuenca b024ee20b8 Merge branch 'onyx' into onyx-vision 2026-08-06 20:38:13 +02:00
Pedro Cuenca eb6046209f Apply norm after token embeddings
This follows the latest transformers approach.
2026-08-06 20:02:12 +02:00
Pedro Cuenca e136efebde Unpermute, to adapt to the latest transformers checkpoint 2026-08-06 19:58:18 +02:00
Pedro Cuenca c784520acb Restore blank line 2026-08-06 14:02:21 +02:00
Pedro Cuenca ebac3dd7ea Small cleanup 2026-08-06 13:57:31 +02:00
Pedro Cuenca 26ed5924e7 No param for rope_theta 2026-08-06 13:52:29 +02:00
Pedro Cuenca df6fe72a7e Patchify via build_inp() 2026-08-06 13:05:32 +02:00
Pedro Cuenca 30d81565cb Make a couple params explicit 2026-08-06 12:52:08 +02:00
Pedro Cuenca 06a8f4eee7 Map to symbolic V_MMPROJ instead of strings 2026-08-06 12:43:04 +02:00
Pedro Cuenca d6e2e02d75 Less params, bilinear pos-emb interpolation as a graph op instead of CPU 2026-08-06 12:10:55 +02:00
Pedro CuencaandYoung Han 6af2931853 Fix token layout
Co-authored-by: Young Han <younghan@fb.com>
2026-08-05 14:35:44 +02:00
Pedro Cuenca 18acf25487 Prefer _size instead of independent _h and _w 2026-08-04 23:39:00 +02:00
Pedro Cuenca 4db53f43bf Additional renames, align with llama.cpp / transformers 2026-08-04 23:31:45 +02:00
Pedro Cuenca 05b8f43f7c Add vision graph
lol, forgot from a previous commit
2026-08-04 23:19:08 +02:00
Pedro Cuenca ef9e715b8c downsample_factor -> merge_size 2026-08-04 23:13:25 +02:00
Pedro Cuenca cd5ba86fe9 Go back to using delimiters.
Otherwise our generations are worse.

Transformers does not use them. We need to trace inputs to verify
whether they are equivalent.
2026-08-04 23:04:06 +02:00
Pedro Cuenca 1ca6e2ab36 Graph 2026-08-04 17:41:40 +02:00
Pedro Cuenca f215b1655d Pre-processing 2026-08-04 17:24:42 +02:00
Pedro Cuenca 68d766eb0d Load mmproj 2026-08-04 16:57:04 +02:00
Pedro Cuenca efed93383d "clip" header declarations 2026-08-04 16:32:54 +02:00
Pedro Cuenca 10bad7ad18 mmproj conversion
Note: some fields to be renamed after the implementation works. We are
keeping compatibility with the reference Meta gguf for testing purposes.
2026-08-04 16:20:24 +02:00
Pedro Cuenca 7c5abdae0f DFlash: inherit rope type from the linked target.
Another option would be to store it in the gguf file itself.
2026-08-03 13:21:29 +02:00
Pedro Cuenca ab29e81d28 Merge branch 'master' into onyx 2026-07-31 19:44:48 +02:00
Pedro Cuenca 38e089b1d6 Register for drafting 2026-07-31 19:42:14 +02:00
Anand Patil 876a432116 vulkan: add POOL_1D op (#25431)
* vulkan : add pool1d push constants and pipeline field

Declared data structures needed for POOL1D OP, which are the vk_op_pool1d_push_constants struct and pipeline_pool1d_f32 field.

* vulkan : add pool1d compute shader

Added pool1d.comp for Vulkan backend mirroring the existing pool2d shader.

* vulkan : add full GGML_OP_POOL_1D support

Added pipeline creation and op dispatch for 1D pooling in the Vulkan backend.

* vulkan : fix pool1d shader logic

Registered pool1d_f32 in vulkan-shaders-gen.cpp and fixed tensor dimension indices and avg pool scale.

* vulkan : fix pool1d end boundary crash and expand test coverage

Fixed an issue where the shader crashed when the end boundary was negative when k0 < p0. Also, added more test cases related to this fix.
b10216
2026-07-31 16:48:58 +02:00
Masato Nakasaka eb41d503ba vulkan: Introduce driver version check for Windows Intel GPU to mitigate crashing (#25192)
* Removed crash guard for Intel

Crash fixed from driver 32.0.101.8860

* Added driver version check for windows

* Change to convert from driverVersion rather than string

* No need to use signed

* Refactor

* allow GPU other than Xe2+

* adjusted function body position
b10215
2026-07-31 16:26:37 +02:00
Xuan-Son NguyenandDaniel Han db7d8b24b5 mtmd: add n_embd_head (#26342)
Co-authored-by: Daniel Han <unslothai@gmail.com>
b10214
2026-07-31 15:30:19 +02:00
Pedro Cuenca e227fc289d No super call; unhardcode eot.
The pattern `self._set_vocab_gpt2()` seems preferred throughout the
codebase, and it allows `set_vocab()` to be called from a different part
of the Python class hierarchy: the drafter model converter that we may
need eventually.
2026-07-31 15:10:06 +02:00
timkhronos a09d8abf8c Support rotated kv cache quant (#26180) b10213 2026-07-31 21:06:40 +08:00
fairydreamingandStanisław Szymczyk 82dbc4f017 llama : load MTP tensors only if they are really used (#26296)
* llama : load MTP tensors only if they are really used

* llama : skip loading MTP (if not used) in remaining models that support MTP

---------

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
b10212
2026-07-31 14:57:02 +02:00
Jeff Bolz 6f3c0a790b vulkan: update vulkan sdk to 1.4.357.0 (#26303) b10211 2026-07-31 07:27:03 -05:00
Pedro Cuenca beebacf8c2 Handle post_norm_eps 2026-07-31 13:45:22 +02:00
Ruixiang WangandGeorgi Gerganov 000547513f server: correct accepted tokens when need draft token replay (#26320)
* spec: correct accepted tokens when need draft token replay

* cont : naming

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b10210
2026-07-31 11:16:17 +03:00
David Friehs 15e755f30d cuda: extract Q2_0 elements via __byte_perm (#25603) b10209 2026-07-31 11:15:44 +03:00
9d9a6d29f6 SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt proc… (#25025)
* SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt processing

* fattn-mkl: fix interleaved dst layout in normalize kernel

- Fix mkl_fa_normalize_head: use interleaved dst layout
  ((query * n_q_heads + head) * DV) matching TILE's
  flash_attn_combine_results. Previously used dense head-major
  layout which wrote head outputs to wrong addresses, corrupting
  attention for all models except Qwen3.6-27B (where GQA=6 heads
  were sparse enough to avoid visible overlap).

- Remove 7 redundant stream->wait() calls — SYCL in-order queue
  already serializes pure SYCL kernel dependencies. Retain only
  the 4 MKL GEMM ↔ SYCL handshake barriers (oneMKL GEMM uses its
  own internal queue that does not respect SYCL in-order).

- Remove unused dst_row_stride, diagnostic clutter, and dead
  K/V hex dump (fa_diag block in fattn-mkl.cpp).

- Add MKL_FA_DISABLE=1 env var for A/B testing.
- Add FA-DISP watchdog (MKL_FA_DEBUG=1) and FA-DIAG output
  fingerprint (MKL_FA_DIAG=1) in fattn.cpp.

Tested: Gemma-4-26B, Gemma-4-31B, Qwen3.6-27B, Qwen3.6-35B-A3B
Perf (B70/Battlemage, 32K, q8_0 KV):
  Gemma-4-26B:  1473 t/s MKL vs 746 TILE (1.97x)
  Qwen3.6-27B:   609 t/s MKL vs 330 TILE (1.85x)

Co-Authored-By: Claude Code on DeepSeek-v4-Pro

* Thank you for the review feedback: rename env vars, use GGML_LOG_INFO, document in SYCL.md

Completed the following:
- Rename MKL_FA_DISABLE → GGML_SYCL_ENABLE_MKL_FA (inverted: 0 to disable)
- Rename MKL_FA_DEBUG → GGML_SYCL_MKL_FA_DEBUG
- Rename MKL_FA_DIAG → GGML_SYCL_MKL_FA_DIAG
- Replace fprintf(stderr, ...) / fflush(stderr) with GGML_LOG_INFO() macro
- Document all three env vars in docs/backend/SYCL.md under Runtime
- Add comment explaining MKL FA activation trigger (flash-attn + quantized
  KV cache + batch-size >= 1024 + n_kv >= 1024)

Resolves review feedback from arthw.
Again, thank you!!!

Co-Authored-By: Claude Code on DeepSeek-v4-Pro

* Thank you for the review feedback round 2: use ggml_sycl_get_env, remove dup waits, gate perf macros

- Replace raw getenv() with ggml_sycl_get_env() in all 4 env-var checks
  (fattn.cpp: GGML_SYCL_ENABLE_MKL_FA, GGML_SYCL_MKL_FA_DEBUG,
   GGML_SYCL_MKL_FA_DIAG; fattn-mkl.cpp: GGML_SYCL_MKL_FA_DEBUG)
- Remove duplicated stream->wait() before ev.wait_and_throw() in GEMM
  KQ and GEMM VKQ — ev.wait_and_throw() already waits for completion
- Gate MKL_ACCUM macro behind do_print so timing accumulators are
  no-ops in normal operation
- Remove redundant MIT/Intel copyright header from fattn-mkl.cpp
- Remove unused #include <cfloat>
- Expand SYCL.md MKL FA docs with step-by-step activation trigger
  and example llama-cli command

Again, thank you!!!

Co-Authored-By: Claude Code on DeepSeek-v4-Pro

* fattn-mkl: enable MKL FA for all KV cache types

Remove the quantized-only restriction on MKL activation — the MKL
kernel converts any non-F16 K/V to F16 via to_fp16_sycl before GEMM,
so F16 (default), BF16, and F32 caches all benefit from XMX hardware
acceleration.  The type restriction was an unnecessary gate.

Before (F16/BF16 default cache + FA on at 32K prefill): ~356 t/s (TILE path)
After:  ~670 t/s (MKL path, matching quantized-cache baseline)

Minimal change: two conditions removed, one comment updated in fattn.cpp.
No kernel or conversion code changes — the dequant pipeline already
covers all types.

* fattn-mkl: rename mkl_disable -> mkl_enable for clarity

* fattn-mkl: refine MKL FA dispatch gates

Three changes:
1. Remove quantized-only restriction - MKL FA activates for all
   KV cache types (F16 default, BF16, F32, quantized).  The MKL
   kernel converts non-F16 K/V via to_fp16_sycl before GEMM.
2. Rename mkl_disable -> mkl_enable to match env var
   (GGML_SYCL_ENABLE_MKL_FA).
3. Replace batch-size threshold with Q->ne[1] >= 32 gate.
   Keeps TG (Q=1) and MTP drafts (Q=3-8) on VEC path where
   fused kernel beats MKL launch overhead.  Routes all
   multi-token prefill through XMX-accelerated GEMM.

Production data confirms Q patterns: 1-8 TG, 32-127 cache reuse,
128+ full reprocess.  At 32K F16/BF16 FA-on: 356 -> 670 t/s.

* ggml-sycl: fix F16 cache + MKL FA multi-turn corruption; add gate guards

Two changes:

1. Always copy F16 K/V to dense row-major buffers before MKL GEMM.
   Previously F16 was read in-place with raw tensor strides. During
   multi-turn conversations, the accumulated KV cache had different
   stride properties than a fresh prefill, producing corrupted outputs.
   Now dense F16 gets a fast memcpy; interleaved (Gemma) gets a strided
   copy kernel. This matches what the quantized paths already did through
   to_fp16_sycl.

2. Gate MKL FA on unsupported op params (max_bias, logit_softcap, batch
   dim mismatch) and pathological F16 strides (nb[1] not a multiple of
   ne[0]*2). These conditions would previously crash inside the MKL
   kernel. Pathological strides (test-only) and ALiBi/softcap fall
   through to TILE/VEC which handle them correctly.

The stride check uses modulo rather than equality, so both dense
(nb1 == ne0*2) and interleaved (nb1 == H * ne0*2) pass — all real
models use these layouts. Only test cases with overlapping rows
(nb1=32 or nb1=75 for ne0=40) are blocked.

Thanks to hmscider for the oneDNN FA PR (#25222) which surfaced the
same insight: always normalize inputs to contiguous F16 before GEMM.

Co-Authored-By: Claude Code using DeepSeek-V4-Pro <noreply@anthropic.com>

* fattn-mkl: fix quant+GQA KV strides, tighten MKL gate, add K>=1024 tests

Adding K>=1024 flash-attn test cases surfaced several MKL bugs:

- Quant K/V with a padded seq-view (real KV cache) used the wrong
  strides in the dequant path... only the true Gemma interleave
  layout should reconstruct strides. nb[2] vs ne[1]*nb[1]
- Gate was firing on shapes the kernel doesn't handle: head_dim < 64
  or not a multiple of 64, MHA, attention sinks, and
  bf16 decode... fell through to vec which no bf16 case.

Gate MKL to the validated envelope: gqa>=2, head_dim 64 through 512
(has to be a multiple of 64) with matching K/V head size, mask,
no sinks/alibi/softcap... everything else falls back to tile.
Covers Qwen Dense/MoE and Gemma4 Dense/MoE

Ran test-backend-ops -o FLASH_ATTN_EXT: 3641/3641 pass.
Perplexity unchanged... 6.7267 MKL vs 6.7290 stock using
Qwen 27b q5_k_xl

* Update ggml/src/ggml-sycl/fattn.cpp

Co-authored-by: Neo Zhang <zhang.jianyu@outlook.com>

* Update ggml/src/ggml-sycl/fattn.cpp

Co-authored-by: Neo Zhang <zhang.jianyu@outlook.com>

* Update ggml/src/ggml-sycl/fattn.cpp

Co-authored-by: Neo Zhang <zhang.jianyu@outlook.com>

* fattn-mkl: bound attention scratch so it doesn't grow with batch or context... also dropped the bf16 comment in fattn.cpp per arthw review.

* Update ggml/src/ggml-sycl/fattn-mkl.cpp

Co-authored-by: Neo Zhang <zhang.jianyu@outlook.com>

* Update ggml/src/ggml-sycl/fattn-mkl.cpp

Co-authored-by: Neo Zhang <zhang.jianyu@outlook.com>

* apply arthw suggestions: enum for dequant modes, macro for wg_size, env-var one-liners

---------

Co-authored-by: Claude Code using DeepSeek-V4-Pro <noreply@anthropic.com>
Co-authored-by: Neo Zhang <zhang.jianyu@outlook.com>
b10208
2026-07-31 10:43:16 +03:00
Neo Zhang d5d3e05bf8 [SYCL] support the missed types in cpy (#26005)
* support the missed types in cpy

* use correct funct

* rm unused code
b10207
2026-07-31 10:25:16 +03:00
fairydreamingandStanisław Szymczyk 69e62fc77c llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized (#25871)
* llama : enforce the same K and V cache types for DeepSeek V4; enable FA if V cache is quantized

* llama : enforce the same K and V cache types for MLA models

---------

Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
b10206
2026-07-31 10:03:30 +03:00
Sachin Sharma 1e22599522 ggml-zendnn : group matmul direct API for mul_mat_id (#25918)
* ggml-zendnn : group matmul API for mul_mat_id

* ggml-zendnn : scale MUL_MAT_ID fallback threshold by expert count
b10205
2026-07-31 09:40:52 +03:00
Neo ZhangandNeo Zhang Jianyu 1c5b89ff63 sycl : support dev2dev memcpy by DEV2DEV_MEMCPY_FORWARD (#26234)
Co-authored-by: Neo Zhang Jianyu <jianyu.zhang@intel.com>
b10204
2026-07-31 09:20:28 +03:00
Neo Zhang a2be61dc87 [SYCL] Support q2 mul_mat (#26231)
* support q2_0 in mul_mat

* support more q2_0 case
b10203
2026-07-31 09:19:41 +03:00
Titaniumtown 1553725965 sycl: fuse RMS_NORM + MUL (#26015) b10202 2026-07-31 09:17:53 +03:00
Masashi Yoshimura 8f4646a63e ggml-webgpu: improve flash_attn_vec for quantized KV at long contexts (#25956)
* improve fa of quantized kv cache

* Fix some bugs and some comments.

* fix v type check and some comments

* Fix build error caused by rebasing

* editorconfig checking pass
b10201
2026-07-31 09:08:40 +03:00
Xuan-Son Nguyen 5f55650a78 mtmd: add lanczos resize method [no release] (#26341) b10200 2026-07-30 21:59:49 +02:00
Xuan-Son Nguyen b4ca032ae3 server: support inp embd to generate next token (#26313)
* server: support embd for sampled token

* fix ~server_batch()
b10199
2026-07-30 21:40:38 +02:00
Jeff Bolz ea63b4d32e vulkan: Support quantized concat (#25684) b10198 2026-07-30 13:11:32 -05:00