Performance work on the MPP tensor API flash attention kernel (NBLK == 2
branch, i.e. dv > 128, e.g. 256/256 prefill). The QK^T + softmax was already
hoisted out of the d-block loop (previous commit); this closes the remaining
gap to the vec kernel:
- C = 32 kv items per chunk (was 64): the smaller K/V operand and P tiles fit
registers better. The pad buffer is written with matching 32-wide chunks
when the tensor path is active (the reserved pad space is 64-wide, a
superset of what the vec kernel needs).
- online softmax with ONE GLOBAL running max (scalar) instead of per-row
maxes: the max cancels in O/S, so any constant >= max score seen so far is
valid; a global max makes the per-chunk O rescale a uniform scalar
multiply with no per-element index decode. The rescale also does not
overlap the tensor core (it serializes between QK^T and PV), so it is
skipped entirely when the max did not change (alpha == 1) - with random or
causal scores the max stabilizes after a few chunks.
- softmax state in per-thread registers (M, S[QPSG], alpha, lmax/lsum), no
shared memory or simdgroup barriers in this branch.
- sinks correction uses the final global max (O and S are both relative to
it); this also removes the per-row max tracking.
Measured on Apple M5 Max (256/256, nq=512, f16 K/V, mask, paired A/B):
kv=20000: vec 25.05 ms, tensor 23.60 ms (1.06x)
kv=10000: vec 8.48 ms, tensor 8.10 ms (1.05x)
(previously ~0.71-0.73x before the hoisting, ~0.82-0.84x after)
test-backend-ops -o FLASH_ATTN_EXT: 4809/4809 pass with the tensor path
enabled and disabled.
Note: 16 queries per threadgroup is 2x faster in isolated microbenchmarks
(K/V-bandwidth bound per query) but spills in the full kernel and is ~30%
slower; 8 queries stays the sweet spot (documented in the kernel).
For DV > 128 (NBLK > 1), the tensor kernel previously ran the entire
QK^T + online-softmax pass once per d block (NBLKx the QK work). This
restructures the NBLK == 2 case (dv = 256) so that QK^T + softmax run
once per KV chunk and the P tile (f32, in registers) is reused for both
PV d blocks (two live PV destination tiles).
NBLK == 1 and NBLK == 4 (dv = 512) keep the per-block form: for NBLK == 4,
carrying 4 live PV destination tiles would overflow the register file.
256/256, nq=512, kv=20000 (f16, mask): 43.4 ms -> 31.1 ms (0.71x -> 0.83x
vs the vec kernel). Full test-backend-ops FLASH_ATTN_EXT suite passes
(tensor on and off).
Adds an optional Flash Attention prefill implementation based on the new
Metal Performance Shaders tensor_ops API (tensor cores, M5).
- kernel_flash_attn_ext_tensor (kernels/fa.metal): two cooperative_tensor
matmuls per KV chunk (QK^T, then PV) in a single-simdgroup threadgroup
(8 queries x 64 KV), f32 online softmax in thread memory, element-wise
output store
- features: causal mask, attention sinks, ALiBi, logit softcap, KV padding
(kvpad), GQA, MLA shapes (dk 64..576, dv 64..512)
- host: runtime feature detection (has_tensor), dispatch gate that falls
back to the existing vec kernel (decode path unchanged), and a
GGML_METAL_FA_TENSOR=0|1 environment override for A/B testing
Verified with test-backend-ops: 4809/4809 FLASH_ATTN_EXT cases pass with
the tensor path enabled (all mask/sinks/ALiBi/softcap/kvpad/GQA
combinations) and disabled (clean fallback).
Note: the tensor_ops API is compiler-managed and context-sensitive (the
register tile layout depends on op, dtypes, opscope and surrounding code).
The kernel layout invariants that keep it correct (f32 P tile, one
simdgroup per threadgroup, element pointer arithmetic on the output) are
documented in the kernel source.
* rpc: support apple RDMA as an RPC transport
* remove set_tensor micro optimization, rpc socket pinning per CR
* remove transparent reconnect
* trigger apple builds on RPC changes
---------
Co-authored-by: Ryan Churaman <rschu@meta.com>
* devops: use GGML_NATIVE=OFF for OpenVINO
Same as in other Dockerfiles.
Should fix#23100
* enable backend dl and cpu all variants
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
* server: fix tool calls getting silently stripped with --prefill-assistant
Last assistant carries tool_calls + --prefill-assistant is on → request
flips into continuation mode, add_generation_prompt forced off, tail
rebuilt from reasoning_content + content only. Tool calls just vanish.
- Auto-continuation now skips trailing assistant msgs that have tool calls
- continue_final_message on those throws a clear error instead of
silently corrupting the prompt
- Regression tests included, red before / green after
Fixes#27588
Developed with AI assistance, disclosed per the contribution policy.
* server : address review: fail on prefill-assistant + trailing tool_calls
Move validation into oaicompat_chat_params_parse (next to the existing
two-or-more-assistant check) and remove it from common_chat_templates_apply,
which has no precedent for validation. Drop the regression tests.
Per review: --prefill-assistant with a trailing assistant message
containing tool calls is not supported and should fail loudly.
* add ccache-buckets action
* use ccache-buckets
* only save on master
* install python3-venv for hip
* add jq and python3 for cuda
* only delete caches older than 5 minutes
* metal : null-check ggml_metal_buffer_init result to avoid OOM crash
ggml_backend_metal_buffer_type_alloc_buffer used the result of
ggml_metal_buffer_init without checking for NULL. ggml_metal_buffer_init
returns NULL when the underlying Metal allocation fails (e.g. an
out-of-memory condition), and the following ggml_metal_buffer_is_shared(res)
call dereferences it, turning a recoverable allocation failure into a hard
crash (EXC_BAD_ACCESS). This is easy to hit on memory-constrained devices
such as iOS when a model/context exceeds the available Metal budget.
Log the failure using the existing GGML_LOG_ERROR convention and return
NULL so the allocator surfaces a diagnosable error up the stack instead of
crashing.
* cont : fix log
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* grammar : accept "\-" escape in character classes
gbnf_escape_char_class() escapes '-' as "\-" but parse_char() rejected
that escape, so generated tool-call grammars failed to parse.
Assisted-by: Claude Code <claude@anthropic.com>
* tests : add parser test for "\-" in char classes
Assisted-by: Claude Code <claude@anthropic.com>
* tests : add integration test for "\-" in char classes
Assisted-by: Claude Code <claude@anthropic.com>
* tests : drop integration and parser tests
* metal : per-device tuned (Q, NE) for flash-attn vec (#25750)
* rebase Q-generic FA vec body from 01dc93607 (#23114)
* add 53 f16 (Q,NE) flash-attn vec instantiations (vec 80 -> 133)
* add FA vec (Q,NE) tuning table + dispatch wiring + SMEM cap fallback
* add FA vec (Q,NE) perf sweep
* fill tuning result
* fold family table into a per-family representative SKU
* refactor tuning result format
* extend FA vec tuning to quantized KV caches
* sync fa vec tuner bucketing with runtime, use pointwise tuning regret
* update tuned table
* format and cleanup
* prefix fa_vec tuning procs with ggml_backend_metal_tuning_, drop unused fa_vec_override_active
* add device id -> token lookup for the offline tuning tool
* add ggml-metal-tuning skeleton
* add op-agnostic perf cell + median timing for the tuner
* add FA-vec graph build + tensor init to the tuner
* tools : add FA-vec (Q,NE) sweep, compression and table emit
* cool down and re-measure the dirty window on thermal drift
* test-backend-ops : replace the FA vec tune mode with a bounded (Q,NE) slice
* tools : document the Metal tuner, point the table comment at it
* abort on unknown KV type, single-source fa_vec_legal_ne
* cleanup
* honor -o in the FA vec (Q,NE) slice
* retune FA-vec (Q, NE) under a pointwise no-harm gate
* cont : add fa-vec tunings for M1 Pro, M2 Ultra, M5 Max
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* metal : per-op source split + parallel compile (#24021)
* preliminary extract common header
* op source split
* split metallib into 8 libs && load in parallel
* derive kernel->library routing from functionNames
* x-macro lib list + underscore filenames, dedup QK_NL, MRC fixes
* op source split 8 to 20
* improve robustness of source fallback
* clean up
* change bool -> atomic_bool
* only prepend headers that source actually includes
* no semaphore, use GCD global queue
* dedup library compile path, fix NSError lifetime, rename gla
* relocate upstream concat/rope_back/repeat kernel changes into split files
* move ggml-common.h from common.h into dequantize.h to shrink binary size
---------
Co-authored-by: lvyichen <lvyichen@stepfun.com>
* metal: add col2im_1d op (f32/f16/bf16) (#25176)
* metal : add set_rows with src0 f16 (#25434)
* metal : add CONV_2D_DW (depthwise convolution) support (#21565)
* metal : add Q2_0 support (#25419)
* metal: fuse snake activation (mul, sin, sqr, mul, add) (#25459)
* ggml-metal: FWHT kernel for metal backend (#25924)
* metal : port new kernels into the split sources
Move the kernels added on master after the split (lightning indexer,
DSv4 hyper-connections, silu_back, f16 bin ops, TQ2_0, the flash-attn KV
dequantization pass, rope offset/inplace, ssm_scan rollback, packed q8_0
dequantization and the tensor-API mat-mat K clamp) into the corresponding
kernels/*.metal sources. Copied verbatim, no functional change.
---------
Co-authored-by: lvyichen <lvyichen@stepfun.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
`repetition_penalty` is standard HF key for repetion penalty.
Currently, only `penalty_repeat` is mapped, read `repetition_penalty`
and map it to `metadata.sampling_penalty_repeat`.
* ci : apply ccache-clear with older/min/dry-run to all ccache jobs
Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
* ci : install gh in ccache-clear if missing (container jobs)
The ccache-clear action relies on the gh CLI, which is not present in
container-based jobs. Install it on demand so those jobs can clear caches.
Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
* ci : install gh via apt repo in ccache-clear
The install.sh script used previously is no longer served (404). Switch to
the official GitHub CLI apt repository, which is still available.
Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
* ci : pass --repo to gh cache commands in ccache-clear
In container jobs gh cannot auto-detect the repository from git, so
gh cache list/delete fail with 'failed to run git: not a git repository'.
Pass the repository explicitly via --repo using GITHUB_REPOSITORY.
Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
* ci : drop -new suffix from vulkan ccache key
The -new suffix was only needed to force a fresh cache. With
ccache-clear now evicting stale caches, the original key can be used
again. The old ccache-vulkan-ubuntu-24.04-arm-new entries still match
the ccache-clear key prefix and are cleaned up automatically.
Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
* ci : fix ccache-clear date parsing on macOS (BSD date)
macOS ships BSD date, which has no -d option. The older cutoff check
was silently disabled there: 'date: illegal option -- d' errors in the
log and the loop was only stopped by the min limit, risking deletion
of caches not older than the cutoff (e.g. saved by a concurrent job).
Parse the ISO-8601 timestamps with GNU date when available and fall
back to BSD date otherwise (TZ=UTC, fractional seconds dropped).
Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
* ci : extract ccache-clear logic into scripts/ccache-clear.sh
The composite action now consists of a dedicated step that installs the
GitHub CLI when missing (e.g. in container jobs) and a thin step that
calls the new script. The script follows the make-release-checks.sh
conventions (usage/env header, set -euo pipefail, CLI flags) and only
checks that gh is available. The action inputs are unchanged, so the
workflow steps are untouched.
Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
* ci : remove unused apple ccaches
* DSV4: sm tensor
* set coarser granularity for head splits
* fix dspark
* add model saving for dsv4 + allow dflash to return on specific device
* add comment about dsv4 seq_rm
* simplify
* add shared expert delayed allreduce
* remove special test for dsv4
* readme : update links
* readme : update maintainer PRs list
Add the new members of the `ggml-org` `maintainers` team to the
author filter of the maintainer PRs link (nikwen, marty1885,
Titaniumtown), keeping the canonical team ordering. The list now
matches the team exactly (35 members).
Assisted-by: pi:llama.cpp/Qwen3.8-27B
Run test-llama-archs with 1 to 4 GGML_METAL_DEVICES, mirroring the
existing CUDA runs, and dispatch the job unconditionally since the
per-backend guards now decide what to run.
Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
The device_info loop iterates over the discovered devices and gets
the available and total memory counts. With the CUDA backend (and
possibly others too) this requires creating a GPU context, which,
in case of CUDA, results in a 550 MB VRAM allocation.
For this information to be used in any way, the log verbosity must
be set to LOG_LEVEL_TRACE. If it's not, including in the default
configuration, the contexts get created, memory sizes get queried,
then the log function quietly discards the data.
In certain cases the user may not want to use any GPU resources.
The device_loop iteration is the only place touching the GPU that
cannot be skipped.
Fix by checking the verbosity level and skipping the loop if there
would be no output.
* DeepseekV4: fix rollback with multi-seq
* fix model loading
* make pending rollback single use
* only clear cache for seq_id for full load
* add assert for compress ratio
* make graph topology static
* pass true instead of flags in clear_compressed
* cont : clean-up + TODOs
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* ui : add browser-style conversation tabs store
Track open conversation tabs in order, persisted to localStorage and
pruned against the loaded conversation list on init. The chat layout
syncs the route's tab on every navigation, so any way of reaching a
conversation opens a tab for it.
* ui : add temporary new-chat tabs
New-chat tabs are unsaved conversations carrying a temporary id used
directly as the route (#/chat/<id>). They live in memory and are only
persisted to the database - keeping the same id so the route and tab
stay stable - when the first message is sent. Deleting one drops it
without confirmation, and deleting conversations now closes their tabs.
* ui : render conversation tab bar in chat layout
Desktop-only tab bar above the chat screen, one tab per open
conversation or new-chat tab. The active tab follows the route id;
clicking navigates, middle-click or the close button closes (switching
to the left neighbor), and a trailing + starts a new chat. Tabs appear
only on chat-id routes; the bare #/ new-chat view has none. The bare
route stays put unless a prompt/model deep-link routes it to a new-chat
tab.
* ui : route new-chat entry points through tabs
The sidebar New chat item, Cmd+Shift+O, the search page and the
arrow-key fallback now open a new-chat tab instead of navigating to the
?new_chat URL, which is removed. New chat is no longer a special route
but a tab like any other conversation.
* ui : track sidebar expanded state in a shared ui store
Move the desktop sidebar expanded/collapsed state out of deviceStore into a
dedicated uiStore so the chat tab bar can react to it.
Assisted-by: pi
* chat : add opt-in conversation tabs setting
Add a Display setting that turns browser-style conversation tabs on or off,
enabled by default.
Assisted-by: pi
* chat : add browser-style conversation tabs with a new-chat screen
Track open conversations as tabs above the chat, one per open chat, plus a
single New chat tab for the bare `#/` route. New chat is just the `#/`
screen - no temporary conversations - and its tab is dropped when navigating
away. Sending the first message creates a real conversation and opens a tab
for it.
Assisted-by: pi
* chat : turn tab bar into a horizontally scrollable carousel
Make the tab bar a horizontally scrollable carousel with edge scroll buttons
and active-tab centering, and align its styling with the sidebar.
Assisted-by: pi
* chat : restyle the scroll-to-bottom button to match tab styling
Assisted-by: pi
* chat : add close-tab keyboard shortcut
Assisted-by: pi
* chat : soften tab bar fade and dim inactive tabs
Assisted-by: pi
* feat: Add stop button to tabs
* refactor: Componentize
* ui : fix carousel scrollability detection
Observe the content wrapper as well as the container, since adding overflowing items does not change the container's own box size. Also expose an onScrollableChange callback.
Assisted-by: pi
* ui : add unified ScrollCarousel component
Single carousel component with top/center variants, gap and scroll options, and hover-revealed chevrons. Rename the HorizontalScrollCarousel accessibility story accordingly.
Assisted-by: pi
* ui : migrate carousels to ScrollCarousel
Switch the settings mobile header, attachments list, thumbnail strip, and MCP resources to the unified component, and drop HorizontalScrollCarousel.
Assisted-by: pi
* ui : improve chat tabs carousel UX
Scroll newly added tabs into view, fade overflowing tabs at the edges, and hide the New chat button while a new-chat tab is open.
Assisted-by: pi
* refactor: Naming
* chat : add keyboard shortcut to jump between conversation tabs
Shift+Cmd/Ctrl+Left/Right cycles the open tabs, mirroring the existing
Shift+Cmd/Ctrl+Up/Down conversation navigation.
Assisted-by: pi
* chat : make the whole tab item act as a link
The full tab is now a link instead of only the inner label button, while
the stop and close buttons stay interactive by swallowing their clicks.
Assisted-by: pi
* chat : adjust tab bar width and use a shared offset variable
Widen the tab bar for the expanded sidebar and rename the tab bar height
variable to --chat-tabs-offset with a smaller value so the chat screen
min-height accounts for the overlay without overshooting.
Assisted-by: pi
* chat : account for the tab bar offset in the assistant min-height
Subtract the tab bar offset when it is shown so the last assistant message
does not overflow the available viewport space.
Assisted-by: pi
* refactor: Post-review fixes
* ui : restore deep links on the chat start page
- handle ?model selection, with ?load=true eager router loading
- ?q now creates a conversation, sends the prompt, and clears the params
- show the not-available-model dialog for unknown models
- never block mount on the conversation list
Assisted-by: pi
* ui : fix tab item link nesting and centralize tab constants
- the tab anchor covers the whole item while stop/close stay siblings,
so interactive elements are never nested inside the anchor
- cmd/ctrl/middle clicks are left to the browser (new window)
- extract the tab labels, the active-tab data attribute, and the
sidebar-offset max widths into constants
Assisted-by: pi
* ui : tidy scroll carousel hook and keep mobile header arrows on
- drop the dead scrollLeft/scrollRight helpers and the unused
onScrollableChange/scrollBy props
- init the carousel once instead of inside a derived
- restore items-start on the center variant
- always show the settings header arrows on touch
Assisted-by: pi
* ui : keep the new-chat tab across reloads and fall back on close
- the new-chat sentinel is no longer pruned on init, so reloading on
the bare new-chat route keeps the tab the user is on
- closing the active conversation falls back to the new-chat screen
when Conversation tabs are off
Assisted-by: pi
* ui : don't block startup on the conversation list
- prune persisted tabs after the list loads in the background instead
of awaiting it during init
- openNewChat now returns void; its return value was never read
Assisted-by: pi
* ui: fix routing nits
* chore: Update doc comments
* refactor: Mark fire-and-forget openNewChat calls as `void`
* chat: fix the deep-linked prompt, the tab width and the tab shortcuts
The chat start page creates the conversation and hands the prompt over
to the chat route, which still sees it in the query string. Sending it
on both sides queues the second copy as a pending message, which shows
up as a stray user bubble once the answer lands and vanishes on reload
since it never reaches the database.
The tab bar takes the max width of the collapsed sidebar while it is
expanded, and the other way round.
The tab list is pruned against a snapshot of the loaded conversations,
so a conversation created while that list is still loading loses its
tab even though the route just opened it. The active tab then falls out
of the list and the cycling shortcut jumps to an edge on every keypress
instead of moving one tab over. Tabs synced from the route are kept as
they are, only the persisted ones are pruned.
The rich chat input claims ctrl or alt with shift and an arrow for its
badge-aware word jump, which now belongs to the tab cycling shortcut.
Holding shift hands the key combination over, the plain word jump is
unchanged.
The close-tab shortcut consumes the event before checking whether the
setting is on, and the logo background loses its importance flag.
---------
Co-authored-by: Pascal <admin@serveurperso.com>