Commit Graph
11262 Commits
Author SHA1 Message Date
Aleksander Grygier a69de44bc3 ui : polish the shell
The secondary button gets its own look back and the chat add button its
own light surface. Pressable elements get a pointer cursor again, the
font rendering smooths on the app shell, the root layout resolves its
props probe through the providers' api url, and the agent skills stay
out of prettier's way.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:27 +02:00
Aleksander Grygier db0eb67ef1 ui : move mcp servers to the sidebar rail
The add menu keeps to what it attaches to the conversation, so the MCP
servers dialog moves to the sidebar rail and its menu entries and the
form callback they used go with it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:27 +02:00
Aleksander Grygier 43dc386afa ui : mount the configuration pane as a drawer in the manager
A row click slides the pane in as a fixed-width drawer: the toolbar's
calls to action leave first, the pane is laid out before its first open,
and a model switch fades one out and the next in. The model information
dialog folds into the pane's information tab, and a chat started from
the pane closes the manager and focuses the composer.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:27 +02:00
Aleksander Grygier 01d8fa0053 ui : add the model configuration pane
The pane shows a model's information, load and inference settings side by
side: what its provider reports, the Hub records it falls back to, the
load controls gated by what the serving backend can do, and the
per-model overrides the load form writes. The slider control arrives
with it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:27 +02:00
Aleksander Grygier 8bc6867ecf ui : Improvements
Adds a measured-px height transition to the tool category groups in the
Tools submenu, replacing the bits-ui Collapsible whose conditional
rendering kills the transition. The same technique as the reasoning panel:
the rows stay mounted, the height animates between 0 and the measured
scrollHeight, and visibility keeps collapsed rows out of the tab order.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 17:12:26 +02:00
Pascal 9fb4766185 ui : resolve bundled icon paths against the app base path
Serve the recommended MCP server favicons and the Hugging Face badge from
the app's base path, not the domain root, so they resolve when the app is
mounted under a path. Also normalizes the safe HTML config indentation.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 17:12:26 +02:00
Aleksander Grygier 7a5d5ecc35 ui : move reasoning effort beside the model selector
The reasoning level becomes its own control next to the selector, and
the add menu drops its reasoning submenu.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:26 +02:00
Aleksander Grygier 156406e99c ui : rework the selector around its providers
The selector keeps working when a provider is down: the trigger shows the
provider's mark and the org avatar, the banner only appears when no
backend is enabled, and the model ids read through their settings. The
model mark becomes MODEL_ICON everywhere.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:26 +02:00
Aleksander Grygier 4d06812a22 ui : group the model selector components
The selector surfaces move under models/ModelsSelector with their own
barrel: the dropdown and sheet relocate, the list, option and trigger icon
split out, and the shared list helpers join the navigation utils. The
searchable dropdown gains a sticky footer and per-surface class hooks.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:26 +02:00
Aleksander Grygier 3897f113a0 ui : show every provider's models in the manager
The table grows a section per enabled provider, each with its own mark,
error and loading state, and a provider filter with live repo counts.
Rows gain their capability gates back, the draft column with its
use-as-draft action, and the quant badge names the provider on its rows.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:26 +02:00
Aleksander Grygier 15860276e6 ui : add the manage providers view
A provider is added from a preset card or by hand: the dialog probes the
endpoint, reads a refusal as a llama.cpp server that wants a key, and
shows the connection test as a status block. Presets carry their official
artwork and a saved backend keeps its branding through the favicon
fallback. The manager dialog gains its third view, and the settings save
stops clobbering keys that other surfaces write.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:26 +02:00
Aleksander Grygier 89c4ba08b2 ui : route chat through the active provider
Chat goes through protocol adapters, so an OpenAI-compatible endpoint
speaks its own wire format: per-backend paths and headers, the model on
the request, tools kept on the local server, and token counts synthesized
for endpoints that do not stream their own timings. The server store keeps
the local server's props while another provider is active, a conversation
resolves the provider its model belongs to before sending, and the chat
screen never blocks on the local probe when the install has none.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:26 +02:00
Aleksander Grygier 88a1c1ba93 ui : add the providers data layer
A provider is a server entry with a base url, an optional key, a protocol
and the paths its API lives at; the local llama.cpp server is the built-in
one. Backends persist in settings and the active one is restored on load.
Requests resolve against the active backend, model ids become
backend-qualified, and every backend's model list is fetched and cached in
the background. The manager's helpers learn to read a model's drafts,
context and the provider that serves it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:26 +02:00
Aleksander Grygier 694f2384cb ui : put the discover view in the models manager
The manager dialog gains its second view: the Discover Models call to
action fades the table out, slides the title with its compass mark, and
fades the discover surface in; the arrow button returns to the table.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:25 +02:00
Aleksander Grygier 3f56d5b38d ui : add the discover models components
The Discover surface: the curated-catalog list with its search and
skeletons, the details pane with its readme, metadata, Hub stats and
download options, and the standalone download progress bar that replaces
the plain one. The quant download button sizes its select with a new xs
trigger, so the select gains that size and the muted box look. The store
stays incomplete when a repo fails, so the next mount retries it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 17:12:25 +02:00
Aleksander Grygier 6226aae1cd ui : add the models manager
A manage models dialog for the local server: a table that folds each repo's
quants into one row, groups them into families, and sorts from its column
headers. Loaded models lead, then favorites, then the local block, then the
hidden one; sections keep their open state in local storage. The toolbar
filters by capability, modality and context, rows act on favorite, delete
and hide, and the manager opens focused on a model from the sidebar or a
download row. The models store gains the recents, hidden and group-open
state, and warms the Hub records for the context and capability columns.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 16:48:49 +02:00
Aleksander Grygier bd03fbcb99 ui : add the shared model row and list primitives
The manager and later surfaces share one set of row parts: the collapsible
section and grouped list containers, the avatar with its org and quant
badges, the capability and context columns with their Hub fallbacks, the
load control, the shared row action set, and the download row with its
progress bar. The toggle and toggle-group controls and the tertiary button
variant arrive with them.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 16:48:49 +02:00
Aleksander Grygier 31220d5c7f ui : render shared model row hints as native titles
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 16:48:48 +02:00
Aleksander Grygier d521981160 ui : remember hub avatars that failed to load
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 16:48:48 +02:00
Aleksander Grygier bf540ab09c ui : shared model display primitives
Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 16:48:48 +02:00
Aleksander Grygier d4e86b6e28 ui : optional sanitized raw HTML in markdown
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 16:48:48 +02:00
Aleksander Grygier 509938e8a9 ui : model download pipeline
Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 16:48:47 +02:00
Aleksander Grygier 5b89ec38f5 ui : model memory-fit estimation
Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 16:48:47 +02:00
Aleksander Grygier 824f42c3e2 ui : strip provider tilde prefix from hub avatar urls
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-09-28 16:48:46 +02:00
Aleksander Grygier ca2d856566 ui : Hugging Face Hub data layer
Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 16:48:46 +02:00
Aleksander Grygier 762e339397 ui : model id grammar for sidecars, quants and capability parsing
Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.

Assisted-by: pi:GLM-5.3-Flash
2026-09-28 16:48:45 +02:00
Aleksander Grygier e54a71f402 ui : type-safe API types, fetch helpers and download-ready models store plumbing
Assisted-by: pi:GLM-5.3-Flash
2026-09-28 16:48:44 +02:00
Adrien Gallouët 6c7a87f7e5 common : fix HF cache paths on Windows (#29475)
Supersedes #29158

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11235
2026-09-28 16:25:24 +02:00
jbooth f00a64c147 webgpu: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor (#29471)
* Fix:  Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor

* Clang formatting
b11234
2026-09-28 16:38:10 +03:00
Georgi Gerganov d77dd0806d tests : refactor test-recurrent-state-rollback (#29426)
* tests : use llama_context_ptr in test-recurrent-state-rollback

Replace raw llama_context pointers with llama_context_ptr and drop the
manual llama_free calls and cleanup lambda.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* tests : run test-recurrent-state-rollback over all dummy models

Add a --models DIR mode that mirrors test-save-load-state: iterate every
dummy model, report PASS/FAIL/SKIP in a table and fail only when a model
fails. Register a single ctest entry with ARGS --models instead of the four
per-model registrations.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* cont : fix typo

* metal : allow fusing 0-element nodes to keep graph packing shape-independent

The fusion packing in ggml_metal_fusion_max excluded 0-element tensors and
the topk_moe/moe_reduce checks rejected n_tokens == 0, so graphs decoding
batches with no outputs packed differently from the worst-case reserved
graph. The Metal optimizer then reordered the nodes differently and
ggml_gallocr_needs_realloc failed on the layout mismatch, forcing an
unexpected graph re-reserve (caught by GGML_SCHED_DEBUG_REALLOC).

Treat empty tensors like their non-empty counterparts: match them in the
pattern sequence and only reject genuinely malformed shapes. Fused kernels
dispatch zero threadgroups for empty graphs, which is a legal no-op.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* tests : run test_multi_seq_split_replay as a separate test

test_multi_seq_split_replay was invoked at the end of test_rollback,
so its result was folded into the rollback status and it only ran when
the rollback part passed.

Give it its own test_status return, run both tests independently over
both cache fills via a shared run_tests helper, and report them as
separate rollback / split replay columns in the --models table with
per-test summaries. The exit code fails when either test fails.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* tests : loosen the split replay nmse bound to 1e-4

test-generate-models seeds its weights from std::random_device, and some
generated lfm2 models drift up to ~1.7e-5 nmse on the split replay due to
rounding noise, tripping the previous 1e-5 bound. Raise the bound to 1e-4
so the random generations stop flaking.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* tests : reuse run_tests_for_model in single-model mode

The single-model path duplicated the model init and the non-recurrent
check from run_tests_for_model; route it through the shared helper
instead. Model load failures now return FAIL rather than SKIP so that
--model with a broken file still exits non-zero, and the helper loads
with model_only like the --models loop does since the tests create
their own contexts.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
b11233
2026-09-28 16:36:38 +03:00
SXX 6f767fe960 ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 (#29423)
* ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86

* add AVX2 support for masked loading and storing in simd_gemm_ukernel_tail

* ggml-cpu: fix FA softcap handling for padded KV tiles
b11232
2026-09-28 16:23:31 +03:00
uvos f916130d00 ci : ignore more vgpr spills in > 256 DQK fattn kernels (#29571) 2026-09-28 14:52:07 +02:00
François-Xavier Gsell 03a667aa30 vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain (#29520) 2026-09-28 14:20:14 +02:00
uvos c2a9e16068 HIP: fix template skip for DKQ > 256 mfma kernels (#29559) b11229 2026-09-28 13:47:29 +02:00
Pascal 4364bf7232 metal: support left and circular padding in GGML_OP_PAD (#29561)
* metal: support left and circular padding in GGML_OP_PAD

Align Metal with CPU, CUDA and Vulkan: shift the source coordinates by
the left paddings, wrap them around with the same wrap_around when
circular, and read the source through nb00, which also fixes a right
padding of a permuted source. A test case covers it.

Drop the f32_4 kernel: its selection is disabled as slower, and it
fails two pad cases once enabled.

* metal: use a function constant for the circular pad variant

Address review from ggerganov: replace the bool template with FC_PAD,
as FC_upscale_aa does, so the pad kernel is compiled once and
specialized per pipeline.
b11228
2026-09-28 12:26:50 +02:00
Sihan YuandGeorgi Gerganov ed7ac35e1e context : do not re-reserve the scheduler when toggling causal_attn (#28751)
* context : do not re-reserve the scheduler when toggling causal_attn

`llama_context::set_causal_attn()` marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive `sched_reserve()` passes per image. This is especially slow for multi-image or video inputs.

The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below).

The re-reserve is unnecessary in this case because `causal_attn` only changes the values written to KQ mask, not tensor shapes or any other buffer sizes.

Note: `causal_attn` is a graph reuse key (`llm_graph_params` via `cparams`), so a new graph is built regardless of `sched_need_reserve`, so this doesn't change the graph rebuilding behaviour.

llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images,
cache_prompt=false, prompt_ms median of 3 (before -> after):

| images | config | H200 before -> after | RTX 4090 before -> after |
|-|-|-|-|
| 1 | `-c 8192 -ub 512` | 134 -> 105 ms (1.27×) | 201 -> 119 ms (1.69×) |
| 24 | `-c 8192 -ub 512` | 2278 -> 1562 ms (1.46×) | 3559 -> 1748 ms (2.04×) |
| 24 | `-c 32768 -ub 2048` | 5379 -> 1584 ms (3.40×) | 13377 -> 1759 ms (7.61×) |

Generated output remains identical before and after.

* qwen4exp : make the indexer bias shape independent of causal_attn

The block/cell bias path was selected on cparams.causal_attn, so the
causal and non-causal graphs differed in tensor shapes and ops. With the
re-reserve removed (previous commit), a runtime flip resulted in
reallocating the compute buffers, which would fail under
GGML_SCHED_NO_REALLOC.

This commit selects the block path from the mask shape only, independent
of causal_attn. causal_attn is instead passed to set_input_qsa.
causal_attn is fixed per graph as it's part of the reuse key. Causal
values are unchanged. Non-causal values now follow the reference rule,
where every visible block competes on score and only unpooled cells are
always selected.

* context : state the causal_attn shape rule in the comment

* cont : add TODOs

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
b11227
2026-09-28 11:58:51 +03:00
Sarah Wu 0c6a6a7ce5 Enables Windows ARM64 build with MSVC cl.exe (#28362)
* can reproduce the issue vlad sees

* fix fma issue

* drop volatile

* fix volatile runtime task

* add arm flag if needed

* fix hsum compile error

* fix syntax in quants

* strengthen sve probing

* make the syntax fixes one liners

* remove debug code

* formatting

* remove macro for float

* drive down gcc instruction count

* support armec

* fix CI comments

address CI comments

fix cross compile issue

remove warning

fix style and fix fma probing

fix style

* add documentation

* update documentation
b11226
2026-09-28 10:07:27 +02:00
Georgi Gerganov 81ef10ea58 tests : fix ggml init (#29554)
* tests : init ggml for test-recurrent-state-rollback

* cont : same for test-save-load-state

* cont : add to test-state-restore-fragmented + add TODOs
b11225
2026-09-28 10:24:21 +03:00
Pascal 5262471615 vulkan: fix wrong results when a mul_mat reads a slice of a larger cache (#28956)
* vulkan: read the batch stride of an in place src0 from nb[2]

A dim01 contiguous tensor can still be a view whose batches are
strided by more than ne[1] rows, the first rows of a KV cache for
example. Both the mat-vec and the matrix paths read such a tensor in
place but passed ne00*ne01 as the batch stride, so every head past
the first read the wrong rows. The same applies to src1. The stride
now comes from nb[2] whenever the tensor is used in place; the value
is unchanged for a contiguous tensor.

test-backend-ops gets an m_v parameter on test_mul_mat, the number of
rows of a in memory, and two cases at the shapes of a decoder self
attention over a cache.

* vulkan: size the in place A and B ranges by their strided extent

The matrix path bound src0 and src1 to the shader with a range of
elements times type size, which ends before the batches of a strided
view. Pipelines with bounded access read zero past that range, so the
same view that the mat-vec path already handles gave wrong results
on Intel and on NVIDIA without coopmat2. The range now comes from
ggml_nbytes when the tensor is read in place.

* vulkan: address review from jeffbolznv

Bind the in place A and B of the matrix path with ggml_vk_subbuffer,
which spans to the end of the buffer, so a strided view is in range
without computing its extent.

mul_mat_id reads the batch stride of an in place src0 and src1 with
the same helper as mul_mat. test_mul_mat_id gets an m_v parameter,
the number of rows of as in memory, and a case whose experts are
strided by more rows than it uses.

* vulkan: read the batch stride of an in place src0 in mul_mat_vec_id

The single token path of mul_mat_id passed ne00*ne01 as the batch
stride of A, so a strided expert view read the wrong rows. The stride
now comes from ggml_vk_batch_stride like the other three paths, and
src1 follows the same rule.

test_mul_mat_id gets a single token case over the strided view.

* vulkan: address review from jeffbolznv

The batch stride of an in place tensor is taken from nb[2] as
nb[2] / type_size * block_size, which holds when nb[2] is padded and
not a multiple of nb[1]. A test_mul_mat case with a padded batch stride
covers it.

* vulkan: keep the A and B ranges exact in mul_mm

The quantized A loads of mul_mm carry no row bound and rely on the
descriptor range to read zeros past the last row of a partial tile.
Binding A and B up to the end of the buffer let those tiles read the
leftovers of a previous node and hung the NVFP4 mul_mm on NVIDIA
without coopmat2. The range is the strided extent of a tensor read in
place and the staged size otherwise.
b11224
2026-09-28 08:36:38 +02:00
Tim Wangandtimothywang21 4da6337767 server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)
* server : allow splitting RANK pooling for causal LLM rerankers

Rerank models fall into two categories: bidirectional cross-encoders
(BERT, etc.) that require all tokens in a single physical batch, and
causal LLMs repurposed as rerankers (Qwen3, Qwen3-VL) that can use
chunked prefill like any other decoder.

Previously the server rejected all RANK-pooling inputs larger than
n_ubatch, and the graph builder hardcoded QWEN3/QWEN3VL arch checks to
determine last-token pooling. This broke long-document and multimodal
reranking for causal models.

Fix: expose llama_get_causal_attn(ctx) so the server can check the
effective runtime attention type (reflecting any --attention override
or set_causal_attn call). Also expose llama_model_is_causal(model)
for querying the static architectural property from GGUF metadata.

can_split() now permits chunked prefill for RANK pooling when the
context is causal. The graph builder's inline arch check is replaced
with the same cparams.causal_attn predicate, removing the duplication.

Assisted-by: Opencode/Qwen3.8-27B

* remove unused llama_model_is_causal, fix whitespace

Assisted-by: opencode

---------

Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
b11223
2026-09-27 23:28:10 +02:00
Georgi Gerganov a97cce86a8 common : avoid side effects around params parsing (#29537)
- register --rpc unconditionally and call llama_supports_rpc() only from its handler
- print server "initialization ..." log after args are parsed

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
b11222
2026-09-27 20:18:56 +03:00
Adrien Gallouët 136887b665 common : make string_split<T> throw on invalid values (#29518)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
b11221
2026-09-27 18:04:11 +02:00
Toki Nasin 9adc7f420c convert : export YaRN scaling parameters for PLaMo-3 (#29528)
Recent PLaMo-3 models use YaRN, while some earlier PLaMo-3 models do not.
The recent PLaMo-3 store their YaRN settings as flat config keys
(rope_scaling_factor, initial_context_length) and build the dict at runtime
in Plamo3Config.rope_parameters. The current converter misses these settings
and writes plain RoPE metadata to GGUF. Mirror the runtime settings into
rope_parameters so the corresponding rope.scaling.* is written to GGUF.
2026-09-27 13:45:47 +02:00
Sigbjørn Skjæret 6fd50a4094 ci : bump ty to 0.0.84 (#29529)
* bump ty to 0.0.84

* fix assertion bug caught by ty
2026-09-27 13:43:08 +02:00
Sigbjørn Skjæret 33c923db1b jinja : add support for dict builtin (#29477)
* add support for dict builtin

* add tests
b11218
2026-09-27 13:41:50 +02:00
lhez c9064dded7 opencl: refine bin kernel loading condition (#29503) b11217 2026-09-27 13:08:39 +03:00
bri-prism c829670992 sycl: FWHT kernels for block widths above 512 (#29243)
The SYCL FWHT covers 64 to 512 via the standard butterfly network, plus
384/640/768/1280 via the Kronecker/Paley construction added separately in
Hadamard hint can produce (1024, 2048, 4096, 8192); those still fall through
to the default case and run as a dense GEMM against the materialized
rotation tensor, correct but O(n^2) instead of O(n log n).

fwht_kernel_wide runs one row per work-group instead of per sub-group, so
each work-item keeps N/NT values rather than N/WARP_SIZE. Butterflies below
the sub-group width still shuffle; those up to the work-group width go
through work-group local memory; the rest stay in registers. Same butterfly
and sign convention as the existing narrow kernel.

ggml's SYCL backend registration (dpct::dev_mgr) unconditionally requires a
GPU-labeled platform to exist and throws before any op-level test can run,
so test-backend-ops could not be exercised on this box (a GPU-less pod) even
via the CPU device. Verified instead with a standalone harness: the same
kernel body run through a real SYCL CPU device (Intel oneAPI DPC++ 2026.1,
OpenCL CPU backend), checked against an independent recursive-doubling
Hadamard reference, cross-validated by first running the existing unmodified
narrow kernel through the identical harness and confirming it passes (rules
out a reference-convention bug before trusting a pass on the new code).
Random-input results for all four widths, single- and multi-row:

  N=1024 NT=256 rows=1  max_abs_err=1.7e-07  max_rel_err=4.9e-04  PASS
  N=2048 NT=256 rows=1  max_abs_err=1.9e-07  max_rel_err=2.0e-04  PASS
  N=4096 NT=256 rows=1  max_abs_err=2.0e-07  max_rel_err=1.4e-04  PASS
  N=8192 NT=256 rows=1  max_abs_err=2.5e-07  max_rel_err=3.8e-03  PASS
  N=1024 NT=256 rows=7  max_abs_err=2.4e-07  max_rel_err=1.0e-03  PASS
  N=2048 NT=256 rows=5  max_abs_err=3.0e-07  max_rel_err=9.4e-04  PASS
  N=4096 NT=256 rows=3  max_abs_err=2.7e-07  max_rel_err=1.7e-03  PASS
  N=8192 NT=256 rows=2  max_abs_err=2.5e-07  max_rel_err=1.9e-03  PASS

This covers the kernel algorithm itself; it does not exercise the ggml
dispatch/supports_op integration end to end, which needs a real GPU (or a
SYCL GPU plugin) to get past backend registration. test-backend-ops build
is verified: fwht.cpp recompiles with zero warnings as part of ggml-sycl.
b11216
2026-09-27 13:08:19 +03:00
Animesh 36d7b08340 CUDA: tune fp16 tile FlashAttention configs for head sizes 40-112 (#26289) b11215 2026-09-27 13:07:37 +03:00
uvos 2ebd9ae621 HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes (#28907)
* HIP: Enable fattn-mma kernel on cdna for dkq > 256 for large batch sizes

* CI: hip-quality-check: ignore spills for very large mfma mma kernels
b11214
2026-09-27 13:06:56 +03:00
Ruben Ortlam cea74625fa vulkan: fix argsort kernel selection for Adreno (#29469) b11213 2026-09-27 13:06:06 +03:00