* ui : shared model display primitives
Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.
Assisted-by: pi:GLM-5.3-Flash
* ui : remember hub avatars that failed to load
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : render shared model row hints as native titles
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : fix badge guard for draft sidecars, keep parameter precision
hasBadges now counts draft sidecar badges, so a sidecar-only model still
renders. Billions keep one decimal for hub counts and stay bare for whole
values. Avatar failures track the org instead of the instance, and the
download progress bar no longer pulses while determinate.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model download pipeline
Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.
Assisted-by: pi:GLM-5.3-Flash
* ui : route sidecar tag lookup through the sidecars util, validate the paused list
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model memory-fit estimation
Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.
Assisted-by: pi:GLM-5.3-Flash
* ui : cover the memory-fit and tool-use heuristics in tests
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : Hugging Face Hub data layer
Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.
Assisted-by: pi:GLM-5.3-Flash
* ui : strip provider tilde prefix from hub avatar urls
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* ui : trim redundant comments in the HF data layer service
Per review: drop JSDoc that restates the method name and inline comments
that restate the code; keep only comments carrying non-obvious context.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : harden the HF data layer error typing, cover the helpers in tests
Carries the HTTP status on retryable fetch errors instead of matching the
message text. Marks expand-dependent catalog fields optional and documents
the data/models index pairing. Adds table tests for the pure helpers.
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : model id grammar for sidecars, quants and capability parsing
Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.
Assisted-by: pi:GLM-5.3-Flash
* ui : escape sidecar tokens in the regex alternation
Assisted-by: pi:zai-org/GLM-5.3-Flash
* ui : type-safe API types, fetch helpers and download-ready models store plumbing
Assisted-by: pi:GLM-5.3-Flash
* ui : document the model list index pairing, fix an em-dash
Assisted-by: pi:zai-org/GLM-5.3-Flash
* Update tools/ui/src/lib/components/app/chat/index.ts
Co-authored-by: Pascal <admin@serveurperso.com>
---------
Co-authored-by: Pascal <admin@serveurperso.com>
* Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx
* Add initial MTP support
* Merge branch optimizations. Reduce allocated compute buffer size, speed up long context decode, fla, and slight MTP improvements.
* Review driven changes, remove env vars, protect tensors
* Strip MTP for initial PR
* Clean up after mtp strip
* Clean up after mtp strip
* Update speculative.cpp
* Update llama-context.h
* Clean up after mtp strip
* Fix tokenizer ignore merges
* Improve quantization protection selection
* Refactor mhc helpers, graph base
* Lint Fixes
* Apply suggestions from code review
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
* Skip glm5-next in model saver, fix CRLF
* Skip glm5-next in sweep
* Remove T4 fallback
* Review cleanup
* Review suggestions
* Defer separate MTP gguf handling to MTP PR, drop filter
* Repad n_head_kv
* kpool init apply
* Order by descending score
* Drop guard
* read kpool from hparams, clarify kpool cache flags, remove kpool_build_state(nullptr)
* Add glm5-next support to model saver and add arch test fixture
* Review cleanup
* Kpool pooled caching clarify
* Add multi stream support
* Finish Rebase
* Sparse FA fir DSA prefill
* Const
* Update llama-model.cpp to fix rebase error
* gguf-py : merge tensor map entries for HC tensors
* model : use build_gdn_l2_norm in GLM5_NEXT implementation
* chore : remove trailing whitespace
* model : use new OP precision setting API in GLM5_NEXT implementation
* mtmd : use ggml_swiglu_clamp in GLM5V and apply the image token limit
The two clamps around swiglu_split are what ggml_swiglu_clamp already does,
so the clamp bounds collapse back to one value. GLM5V also never called
set_limit_image_tokens(), so --image-max-tokens had no effect.
Assisted-by: Claude Opus 5
(cherry picked from commit 46d18e12d422be4cc04a70e4a9a9e0168bb3d5b7)
* llama : keep the GLM5-Next k-pool layout across ubatches
The layout was rebuilt from a full cell scan on every ubatch. Pools are fixed
by the positions relative to the sequence's first one, so the layout now lives
on the memory and a ubatch only appends to it.
A sequence edit no longer stales every pooled key either, only the ones at or
after the edited position, which makes a tail seq_rm free. The pooling subgraph
is built unconditionally so the graph shape no longer changes every kpool
tokens, and the pool axis is folded into rows before soft_max, which otherwise
exceeds the CUDA gridDim.y limit past n_kv 262144.
Assisted-by: Claude Opus 5
(cherry picked from commit 5d1c40b93e17fddbf73b785efe43e0d02ccb3977)
* model : write the GLM5-Next recurrent rollback checkpoints
The conv state and the delta net state were only written to the live row, so a
rollback restored whatever the checkpoint rows happened to hold. Take the same
route as kimi-k3: build_recurrent_attn for the state, and write all K_rs conv
groups. That also drops a state view that assumed contiguous rows.
Enroll the arch in test-recurrent-state-rollback, which catches this under its
garbage-filled cache pass.
Assisted-by: Claude Opus 5
(cherry picked from commit 5ace37e86d5d448e83ef5dde5632c748185b18cd)
* llama: fix PR #27773 test-save-load-state restore failure
Clear the attention and indexer cache data after a failed hybrid state restore so restored NaNs cannot affect a later sequence.
Assisted-by: Codex
* llama: fix PR #27773 gpu-rocm graph reallocation
Reserve the full GLM5-Next pool capacity and dirty pool count. The gpu-rocm Test step aborts when n_new grows while the graph node count stays fixed; CUDA, Vulkan, Metal, and WebGPU checks report the same error.
Assisted-by: Codex
* llama : fix GLM5-Next k-pool layout staleness after edits and shared teardown
Two defects in the cross-ubatch k-pool layout added by the k-pool commit:
1. Wrong results. An edited sequence only rebuilt its pool layout when its cell
count changed, so if the first ubatch after an edit added back exactly as many
cells as were removed, the stale position-to-cell list survived. With a unified
cache and more than one sequence, where another sequence takes the freed cells,
the reused layout points at the wrong cells (CPU: large logit drift, CUDA: NaN).
Rebuild whenever the sequence is stale, not only on a size mismatch.
2. Slowdown. "shared" mode was assumed to end only with an edit that forces a
rebuild, but sharing also ends when the other sequence is removed. The survivor
kept shared = true, pinning cache_safe off and re-pooling every pool on every
ubatch (server trigger: n>1 completions with -kvu, via the seq_cp in
copy_state_to). In seq_rm, if the layout has shared cells, stale every sequence
so one rebuild re-derives sharing and cache_safe returns to 1.
Assisted-by: Claude Opus 5
* llama : fix build_attn_mha stream stride for non-contiguous q
build_attn_mha split the batch into streams with a stream stride of
q->nb[3]/n_stream. That only equals one stream's span, (ne[2]/n_stream)*nb[2],
when q is contiguous. GLM5-Next is nope-only, so it does not concat a rope part
and passes the permuted q_absorbed straight in, where nb[3] != ne[2]*nb[2]; the
stride was then n_head times too large and every stream s >= 1 read another
head's queries. Split-KV (-np N without --kv-unified) multi-stream prefill was
wrong for every stream past the first. Unified KV and decode were unaffected
(n_stream == 1, and decode takes the gather path). Other MLA models concat rope
so q is contiguous and the computed value is unchanged for them.
Compute the stride from the token dimension, which is identical for a
contiguous q.
Assisted-by: Claude Opus 5
* llama : re-derive GLM5-Next k-pool sharing on state_read/state_drop
The shared-cell teardown added to seq_rm (stale every sequence when the layout
has shared cells, so a survivor does not keep shared = true and pin cache_safe
off) was missing from the other paths that can free shared cells: state_read
and state_drop staled only the one sequence. Apply the same re-derivation there
and correct the comment that claimed sharing ends only via an edit or seq_rm.
Assisted-by: Claude Opus 5
* quant : drop duplicate GLM5-Next hc_ filter
The hc_ name filter was listed twice in the GLM5_NEXT protection block.
Assisted-by: Claude Opus 5
* glm5-next: scope K-pool cache access to indexed operations
* glm5-next: keep K-pool access in hybrid index memory
* glm5-next: keep mHC graph builders model-local
* glm5-next: mark only touched pools per ubatch
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Piotr Wilkin <ilintar@gmail.com>
With --path or --no-ui, /sw.js returned 404, and a 404 does not remove a service worker, so browsers kept showing the cached built-in UI. Serve a worker that unregisters itself, clears its caches and reloads open tabs. A sw.js in the --path folder is still served first.
Assisted-by: Claude Opus 5.5
* server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding)
Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.
The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.
Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.
Assisted-by: Opencode Qwen3.8 27B
* clean up comments and docs
* refactor
* add tests
* support video and audio inp
---------
Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
* models: pad on the left with ggml_pad_ext
The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders
build a left padding as a right pad followed by a roll, and DFlash2
concatenates a zero filled block in front of the previous tokens.
ggml_pad_ext does both in one node now that every backend supports a
left padding. The Gemma 4 audio embeddings are bit identical.
* models: skip the DFlash2 taps that only read padding
A tap at or past block_size shifts every row out of the block, so its
term is zero. The loop runs min(kernel_size, block_size) taps.
* adapt common
* add common_batch
* wip
* wip: spec
* cont
* common_speculative_process
* server_batch to use common_batch
* rm some stale calls
Assisted-by: Claude Fable 5.1
* migrate mtmd
* handle imrope, handle return val of add()/add_embd()
* add spec zeros vector
* add warning on zero fill path
* server : allow splitting RANK pooling for causal LLM rerankers
Rerank models fall into two categories: bidirectional cross-encoders
(BERT, etc.) that require all tokens in a single physical batch, and
causal LLMs repurposed as rerankers (Qwen3, Qwen3-VL) that can use
chunked prefill like any other decoder.
Previously the server rejected all RANK-pooling inputs larger than
n_ubatch, and the graph builder hardcoded QWEN3/QWEN3VL arch checks to
determine last-token pooling. This broke long-document and multimodal
reranking for causal models.
Fix: expose llama_get_causal_attn(ctx) so the server can check the
effective runtime attention type (reflecting any --attention override
or set_causal_attn call). Also expose llama_model_is_causal(model)
for querying the static architectural property from GGUF metadata.
can_split() now permits chunked prefill for RANK pooling when the
context is causal. The graph builder's inline arch check is replaced
with the same cparams.causal_attn predicate, removing the duplication.
Assisted-by: Opencode/Qwen3.8-27B
* remove unused llama_model_is_causal, fix whitespace
Assisted-by: opencode
---------
Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
- register --rpc unconditionally and call llama_supports_rpc() only from its handler
- print server "initialization ..." log after args are parsed
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
The original function was broken on Windows for some unicode paths
Paths without a trailing separator now create the last directory too,
matching the function name. All current callers already include a
trailing separator, so this change does not affect them.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese
test utterances. In Japanese, some differences changed entire words.
This change:
* uses `log(x + 2^-24)` instead of clamping to the log floor
* uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)`
* adds the normalization epsilon to the standard deviation instead of inside the square root
Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged.
Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against
http://github.com/Liquid4All/liquid-audio fp32.
Test set:
* 200 LibriSpeech `test-clean` utterances (EN)
* 200 Common Voice `ja` test utterances (JP)
* identical 16 kHz audio passed to both implementations
| Greedy transcript identical to `liquid-audio` | Without fix | With fix |
| --------------------------------------------- | ----------: | ----------: |
| EN F16 | 191/200 | 200/200 |
| JP F32 | 187/200 | 200/200 |
| JP F16 | 187/200 | 199/200 |
The remaining JP F16 difference is a comma and matches the reference implementation's own bf16
output.
Mel relative L2 error versus `liquid-audio`:
* EN: 3.2% -> ~2e-6 median
* JP: 3.9% -> ~2e-6 median
* model : fold Ling 3.0 VL into the BailingMoeV3 architecture
Assisted-by: Scout
* model : keep shared NORM rope list intact when gating bailingmoe3 on mrope sections
---------
Co-authored-by: aetherbird <aetherbird@users.noreply.github.com>
The OpenAI chat completions API specifies content part type "video_url"
with a {"url": ...} object, and clients typically send data: URIs
(e.g. data:video/mp4;base64,...). The llama-server only accepted the
non-standard "input_video" type and rejected data: URIs for video
(accept_base64_uri=false), so any OpenAI-conformant client failed with
"unsupported content[].type" or "Invalid uri format".
- accept "video_url" as an alias of "input_video"
- read the media object from whichever key was used
- allow data: URIs for video (data:video/*), as already done for images
* server: route every model load through the queue
A model loaded by the fast path has no queue entry, so tick() evicts
it at its LOADED transition before its own request is proxied. Every
load now joins the queue, whose entry protects the model until its
waiters leave.
* server: do not admit requests into a stopping model
A request for a model that is being stopped still sees it LOADED and
is proxied into the dying child. Such a request now joins the queue
and is served by the next instance. The stopping mark is cleared
under the same lock that sets UNLOADED, so no request can see a
model that is neither stopping nor unloaded while its child is gone.
This commit tweaks the Toaster element to include a close button.
These toasts often cover other UI elements like the model selector, and
this change avoids having to wait for them to disappear on their own
(e.g. after a load failure).
Allow configuring --temp, --top-p, --min-p, --repeat-penalty,
--presence-penalty and --frequency-penalty via LLAMA_ARG_* so
llama-server can be fully controlled from an EnvironmentFile
(e.g. systemd on Debian).
Use `llama-gen-docs` to regenerate the readme files.
In router mode, authentication belongs to the router. unset_reserved_args()
already unset LLAMA_API_KEY, but did not unset LLAMA_ARG_API_KEY_FILE.
When --api-key-file was passed, children re-validated against file keys only,
causing clients using --api-key to 401 on chat completions (#28820).
In addition, router internal calls without auth headers (such as
POST /v1/streams/lookup and DELETE /v1/stream) were silently rejected with 401.
Unset LLAMA_ARG_API_KEY_FILE in unset_reserved_args() so no API keys reach
child instances. This keeps keys out of child argv, ensures all keys the router
accepts work end-to-end, and prevents router internal stream calls from 401ing.
Fixes#28820
* ui : let the chat column shrink below its content width
The chat column is a flex item, so its automatic minimum size kept it as wide as the widest row inside it. Message rows cap at max-w-3xl plus padding, so a narrower window pushed a page-level horizontal scrollbar.
Set min-w-0 on the column so the inner scroll containers take over.
Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash
* ui : wrap markdown tables in a scroll container
Markdown tables render as a bare <table>, which keeps its content-driven minimum width and can stretch the chat column past the window. The table-wrapper CSS already existed, but nothing produced the wrapper.
Add a rehype plugin that wraps each table in div.table-wrapper, following the existing enhance-* plugins.
Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash
* ui : scroll long inline content inside markdown blocks
Long unbreakable content (inline code, paths, hashes) widened the message row and spilled over the neighbour elements. Give each markdown block a horizontal scroll container, and the content root one as well, since the trailing block renders with display: contents and has no box of its own.
Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash
* ui : use exact transition properties for markdown images
transition: all repainted every property and 300ms felt sluggish. Name transform and box-shadow at 200ms ease-out, and gate the hover scale behind (hover: hover) and (pointer: fine) so touch taps do not trigger it.
Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash
* ui : fit wide image attachments to the message width
Attachment thumbnails used a fixed height with w-auto, so a wide image kept its aspect-driven width and, being flex-shrink-0 in a right-aligned bubble, overflowed to the left of the message row.
Cap the thumbnail with max-height and max-width instead of a fixed height so it scales down proportionally, and let it shrink outside the single-row carousel.
Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash
* ui : keep long tool call titles inside the message row
A tool title could not shrink below its content, so a long path escaped the message row. Let the title span shrink and scroll, and for the file tools put the value on its own line only when it does not fit, with the value as the only scroll container.
Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash
* ui : render get info as a collapsible block with a table
get_info rendered its own always-open row with the values trailing the label. Use the shared ToolCallBlock chrome so it collapses like the other tools, and list os and cwd as table rows with the key as a row header.
The error and pending states now show inside the body, including the plain-string errors the server tools path produces.
Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash
* test : pin the server mode in the add menu a11y story
The story asserts the add menu's first enabled item is the reasoning submenu, which is mounted only outside router mode. The vitest dev server proxies /props to whichever server is running, so the assertion depended on the machine's server mode and failed whenever a router was up.
Pin the mode in the story, including props.role so a re-detection cannot flip it back.
Assisted-by: pi:deepseek-ai/DeepSeek-V4.1-Flash
* ui: wrap long markdown tokens instead of scrolling every block
Making each markdown block and the content root a horizontal scroll
container turns any hover transform into a scrollbar: the blockquote
translate and the image zoom overflow their block and flash a scrollbar
under it. Each block also becomes a block formatting context, so the
paragraph margins stop collapsing across blocks and the spacing doubles.
Drop both overflow-x rules and let long unbreakable tokens wrap with
overflow-wrap: break-word on the content root. break-word leaves the
min-content width untouched, so wide tables and code blocks keep
scrolling inside their own containers.
---------
Co-authored-by: Pascal <admin@serveurperso.com>
* server-models : show source per model in log
- Show [source] tag (preset/models_dir/cache) per model instead of cryptic * marker
- Show HF hub cache path in the 'Loaded cached model presets' log
- Add hf_cache::get_cache_dir() public accessor
Assisted-by: pi:llama.cpp/Qwen3.8-27B
* cont : pad log
* ui: fix accidentally removed reasoning menu in single model mode on desktop
* ui: formatting task run to fix storybook test
* ui: mount the add menu reasoning submenu outside router mode only
The models selector already owns the reasoning submenu in router mode,
so the add menu only mounts it in single model mode. The first enabled
item of the add menu is now the reasoning submenu, the accessibility
story expects it.
---------
Co-authored-by: Ben Babik <work@benjaminbabik.com>
Co-authored-by: Pascal <admin@serveurperso.com>