Commit Graph
1371 Commits
Author SHA1 Message Date
Aleksander Grygier 486a7bd06d ui : show every provider's models in the manager
The table grows a section per enabled provider, each with its own mark,
error and loading state, and a provider filter with live repo counts.
Rows gain their capability gates back, the draft column with its
use-as-draft action, and the quant badge names the provider on its rows.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier cc229b4b6d ui : add the manage providers view
A provider is added from a preset card or by hand: the dialog probes the
endpoint, reads a refusal as a llama.cpp server that wants a key, and
shows the connection test as a status block. Presets carry their official
artwork and a saved backend keeps its branding through the favicon
fallback. The manager dialog gains its third view, and the settings save
stops clobbering keys that other surfaces write.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 742fe79906 ui : route chat through the active provider
Chat goes through protocol adapters, so an OpenAI-compatible endpoint
speaks its own wire format: per-backend paths and headers, the model on
the request, tools kept on the local server, and token counts synthesized
for endpoints that do not stream their own timings. The server store keeps
the local server's props while another provider is active, a conversation
resolves the provider its model belongs to before sending, and the chat
screen never blocks on the local probe when the install has none.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 51843910ef ui : add the providers data layer
A provider is a server entry with a base url, an optional key, a protocol
and the paths its API lives at; the local llama.cpp server is the built-in
one. Backends persist in settings and the active one is restored on load.
Requests resolve against the active backend, model ids become
backend-qualified, and every backend's model list is fetched and cached in
the background. The manager's helpers learn to read a model's drafts,
context and the provider that serves it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 489f6e1f09 ui : show a paused download's state on its chip
A paused quant chip reads as paused at rest, the way a running one shows
its spinner, and offers the resume action on hover. The selector's
download row gets a View in Discover action beside its cancel.

Assisted-by: pi
2026-10-02 11:59:25 +02:00
Aleksander Grygier 7cec6ff958 ui : remember unfinished model downloads
A partial file is named after its content hash, so nothing on disk says
which model it belongs to once the router restarts. The router now keeps
the tags of the downloads it started, clears a tag when the download
finishes or is deleted, and lists what is left as still downloading
without starting anything. The UI adopts those tags as paused on every
list fetch, so a refresh or a fresh open shows an incomplete download and
can resume it from the files already on disk.

Assisted-by: pi
2026-10-02 11:59:25 +02:00
Aleksander Grygier ae433124b3 ui : make resuming a paused download work
Re-posting the tag of a paused download hit the router's "already
exists" check, because a paused download parks its entry as DOWNLOADED
until the next reload: a download entry that is not running is now
replaced and resumed. The UI keeps the paused row when the request still
fails, instead of dropping it and looking finished.

Assisted-by: pi
2026-10-02 11:59:25 +02:00
Aleksander Grygier 38c2c9cbc2 ui : lead the discover list with downloads
The discover sidebar lists running downloads above the suggested models,
and steps aside while a query is active. A paused download stays in the
manager table, where it is resumed or dropped.

Assisted-by: pi
2026-10-02 11:59:25 +02:00
Aleksander Grygier 1120c39205 ui : resume model downloads from the pause point
A stopped download keeps its partial file instead of deleting it, so
re-posting the tag resumes from disk. The first progress record now
carries the resume offset, the paused snapshot survives into the
resumed entry, and a pause settles as soon as it is requested.

Assisted-by: pi
2026-10-02 11:59:25 +02:00
Aleksander Grygier 06b21f7bba ui : add the Enable Discover Models setting
The flag gates browsing and downloading HuggingFace GGUF models, so it is
declared on the discover layer; the repo-tree lookup in the manager layer
above reads it from here.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 4881f73d31 ui : put the discover view in the models manager
The manager dialog gains its second view: the Discover Models call to
action fades the table out, slides the title with its compass mark, and
fades the discover surface in; the arrow button returns to the table.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 452d055f91 ui : add the discover models components
The Discover surface: the curated-catalog list with its search and
skeletons, the details pane with its readme, metadata, Hub stats and
download options, and the standalone download progress bar that replaces
the plain one. The quant download button sizes its select with a new xs
trigger, so the select gains that size and the muted box look. The store
stays incomplete when a repo fails, so the next mount retries it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier 796fbb571f ui : optional sanitized raw HTML in markdown
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.

Assisted-by: pi:GLM-5.3-Flash
2026-10-02 11:59:25 +02:00
Aleksander Grygier ee5685f8cd ui : refine the model pane and the manager filters
The pane no longer says the server withholds data: it says the metadata
and the chat template appear once the model is loaded, and the template
block grows with its content instead of scrolling inside a fixed box. The
pane ends with a Delete this model from disk action, confirmed by the same
dialog the table row uses and hovering like a destructive dropdown entry.
The filter toggles list only the modalities the models at hand actually
carry, and starting a new chat from the pane closes the dialog.

Assisted-by: pi
2026-10-02 11:59:24 +02:00
Aleksander Grygier 7bc505dab7 ui : list downloads as their own manager section
The models table lists in-flight and paused downloads in a Downloading
section: flat rows like Loaded models, right after them, with a full
width progress bar along the bottom. The status column shows the percent
and swaps in pause or resume on hover; pause, resume and delete live in
the row's action menu. The selector lists running downloads only.

Assisted-by: pi
2026-10-02 11:59:24 +02:00
Aleksander Grygier 7851efc3b2 ui : default to the Hub metadata and hide avatars when it is off
Turn the Hugging Face Hub models-metadata setting on by default, and when
it is off render no org avatar at all instead of the monogram fallback:
ModelOrgAvatar is the choke point, ModelAvatar stops mounting its wrapper,
and the base-model lookups in the selector and the download row are skipped.

Drop the vestigial "Enable Discover Models" setting from this branch: the
discover UI lives on the discover branch, and nothing here referenced it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier 0f18f69694 ui : move mcp servers to the sidebar rail
The add menu keeps to what it attaches to the conversation, so the MCP
servers dialog moves to the sidebar rail and its menu entries and the
form callback they used go with it.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier 93732021e7 ui : drop an orphaned doc comment left by a moved method
Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier cfa83f4a1b ui : move reasoning effort beside the model selector
The reasoning level becomes its own control next to the selector, and
the add menu drops its reasoning submenu.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier ca4cc8a541 ui : rework the selector around its providers
The selector keeps working when a provider is down: the trigger shows the
provider's mark and the org avatar, the banner only appears when no
backend is enabled, and the model ids read through their settings. The
model mark becomes MODEL_ICON everywhere.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier b0d0ee6b50 ui : group the model selector components
The selector surfaces move under models/ModelsSelector with their own
barrel: the dropdown and sheet relocate, the list, option and trigger icon
split out, and the shared list helpers join the navigation utils. The
searchable dropdown gains a sticky footer and per-surface class hooks.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Aleksander Grygier 5fe9a1ad67 ui : add the models manager
A manage models dialog for the local server: a table that folds each repo's
quants into one row, groups them into families, and sorts from its column
headers. Loaded models lead, then favorites, then the local block, then the
hidden one; sections keep their open state in local storage. The toolbar
filters by capability, modality and context, rows act on favorite, delete
and hide, and the manager opens focused on a model from the sidebar or a
download row. The models store gains the recents, hidden and group-open
state, and warms the Hub records for the context and capability columns.

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
2026-10-02 11:59:24 +02:00
Xuan-Son Nguyen a4cb4c61fd llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) (#29818)
* init conversion

* convert: ok

* model loaded

* add server code

* improve conversion script

* support shared prompt prefix

* add docs, imorove UX a bit

* add vision support

* add openjev tiny model for testing

* add dev docs

* support lev & kev

* clean up

* fix lev noul

* fix py lint

* nits docs

* clarify about not supporting date_facts
2026-10-02 11:56:04 +02:00
Adrien Gallouët 70849ee82c common : remove fs_open_ifstream() by using u8path() (#29841)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-10-02 11:43:45 +02:00
Adrien Gallouët f1cee9941b common,rpc : fix cache dir creation through symlinks on buggy libstdc++ (#29816)
See https://gcc.gnu.org/bugzilla/show_bug.cgi?id=101510

Close #29759

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-10-01 20:15:02 +02:00
Xuan-Son Nguyen 68e79bd8cd skill: note about model-specific CLI arguments + testings (#29808)
* skill: note about adding model-specific CLI arguments

* add testing instructions
2026-10-01 19:53:23 +02:00
Sam Malayek d775ebf363 server: return HTTP 400 for invalid embedding requests (#29060) 2026-10-01 16:52:43 +02:00
Xuan-Son Nguyen 552f18f912 mtmd: cap max_image to ubatch for non_causal models (#29773) 2026-10-01 11:55:11 +02:00
Kushal Garg 7dad6db858 llama-bench : fix verbosity filter to show GGML_LOG_ERROR (#28229)
* bench : fix verbosity filter to show GGML_LOG_ERROR (#28107)

* bench: remove dead variables
2026-10-01 08:20:00 +03:00
Marlon Paz 79625e056e llama-bench : fix docs (#29464)
* OoD documenatation for llama-bench

Signed-off-by: mairp <oec.valle.art@gmail.com>

* Unset default: auto

Signed-off-by: mairp <oec.valle.art@gmail.com>

---------

Signed-off-by: mairp <oec.valle.art@gmail.com>
2026-10-01 08:18:46 +03:00
Xuan-Son Nguyen 60e9cf7a7b batch: migrate the rest of examples to llama_batch_ext (#29601)
* migrate the rest

* test-thread-safety

* rm common_batch_staged
2026-09-30 18:08:43 +02:00
Pascal b04642061d cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast (#29722)
* cli: exit on stdin EOF and drop the console wide Ctrl+C broadcast

On Windows the simple input reader sends CTRL_C_EVENT to every process
attached to the console when stdin reaches EOF, killing unrelated
processes such as a supervising agent. The CLI only stopped on EOF
because of that self inflicted SIGINT; on POSIX, and with the advanced
reader, it spins forever printing prompts.

Drop the broadcast so both platforms just return an empty read, and
treat an empty read as EOF in the chat loop and the model selection,
since a submitted line always ends with a newline.

* cli: keep the newline of a trailing "/" and stop mtmd-cli on EOF

A lone "/" came back as an empty read and was taken for EOF, and
mtmd-cli only stopped on EOF through the removed broadcast.
2026-09-30 16:24:46 +02:00
Aleksander Grygier 4a096b8ff6 ui : shared model display primitives (#29644)
* ui : shared model display primitives

Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.

Assisted-by: pi:GLM-5.3-Flash

* ui : remember hub avatars that failed to load

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : render shared model row hints as native titles

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : fix badge guard for draft sidecars, keep parameter precision

hasBadges now counts draft sidecar badges, so a sidecar-only model still
renders. Billions keep one decimal for hub counts and stay bare for whole
values. Avatar failures track the org instead of the instance, and the
download progress bar no longer pulses while determinate.

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:18 +03:00
Aleksander Grygier 8664eaea30 ui : model download pipeline (#27959)
* ui : model download pipeline

Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.

Assisted-by: pi:GLM-5.3-Flash

* ui : route sidecar tag lookup through the sidecars util, validate the paused list

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:17 +03:00
Aleksander Grygier 4cfb6d1c75 ui : model memory-fit estimation (#27957)
* ui : model memory-fit estimation

Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.

Assisted-by: pi:GLM-5.3-Flash

* ui : cover the memory-fit and tool-use heuristics in tests

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:17 +03:00
Aleksander Grygier 9b43336114 ui : Hugging Face Hub data layer (#27947)
* ui : Hugging Face Hub data layer

Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.

Assisted-by: pi:GLM-5.3-Flash

* ui : strip provider tilde prefix from hub avatar urls

Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash

* ui : trim redundant comments in the HF data layer service

Per review: drop JSDoc that restates the method name and inline comments
that restate the code; keep only comments carrying non-obvious context.

Assisted-by: pi:zai-org/GLM-5.3-Flash

* ui : harden the HF data layer error typing, cover the helpers in tests

Carries the HTTP status on retryable fetch errors instead of matching the
message text. Marks expand-dependent catalog fields optional and documents
the data/models index pairing. Adds table tests for the pure helpers.

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:16 +03:00
Aleksander Grygier f653250407 ui : model id grammar for sidecars, quants and capability parsing (#27946)
* ui : model id grammar for sidecars, quants and capability parsing

Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.

Assisted-by: pi:GLM-5.3-Flash

* ui : escape sidecar tokens in the regex alternation

Assisted-by: pi:zai-org/GLM-5.3-Flash
2026-09-30 13:42:16 +03:00
Aleksander GrygierandPascal fa2bde5543 ui : type-safe API types, fetch helpers and download-ready models store plumbing (#29582)
* ui : type-safe API types, fetch helpers and download-ready models store plumbing

Assisted-by: pi:GLM-5.3-Flash

* ui : document the model list index pairing, fix an em-dash

Assisted-by: pi:zai-org/GLM-5.3-Flash

* Update tools/ui/src/lib/components/app/chat/index.ts

Co-authored-by: Pascal <admin@serveurperso.com>

---------

Co-authored-by: Pascal <admin@serveurperso.com>
2026-09-30 13:42:15 +03:00
649dcb1036 add GLM-5.3-Flash (GLM5-Next) support (#27773)
* Rebase GLM-Next support onto master, and migrate to llama-memory-hybrid-idx

* Add initial MTP support

* Merge branch optimizations. Reduce allocated compute buffer size, speed up long context decode, fla, and slight MTP improvements.

* Review driven changes, remove env vars, protect tensors

* Strip MTP for initial PR

* Clean up after mtp strip

* Clean up after mtp strip

* Update speculative.cpp

* Update llama-context.h

* Clean up after mtp strip

* Fix tokenizer ignore merges

* Improve quantization protection selection

* Refactor mhc helpers, graph base

* Lint Fixes

* Apply suggestions from code review

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>

* Skip glm5-next in model saver, fix CRLF

* Skip glm5-next in sweep

* Remove T4 fallback

* Review cleanup

* Review suggestions

* Defer separate MTP gguf handling to MTP PR, drop filter

* Repad n_head_kv

* kpool init apply

* Order by descending score

* Drop guard

* read kpool from hparams, clarify kpool cache flags, remove kpool_build_state(nullptr)

* Add glm5-next support to model saver and add arch test fixture

* Review cleanup

* Kpool pooled caching clarify

* Add multi stream support

* Finish Rebase

* Sparse FA fir DSA prefill

* Const

* Update llama-model.cpp to fix rebase error

* gguf-py : merge tensor map entries for HC tensors

* model : use build_gdn_l2_norm in GLM5_NEXT implementation

* chore : remove trailing whitespace

* model : use new OP precision setting API in GLM5_NEXT implementation

* mtmd : use ggml_swiglu_clamp in GLM5V and apply the image token limit

The two clamps around swiglu_split are what ggml_swiglu_clamp already does,
so the clamp bounds collapse back to one value. GLM5V also never called
set_limit_image_tokens(), so --image-max-tokens had no effect.

Assisted-by: Claude Opus 5
(cherry picked from commit 46d18e12d422be4cc04a70e4a9a9e0168bb3d5b7)

* llama : keep the GLM5-Next k-pool layout across ubatches

The layout was rebuilt from a full cell scan on every ubatch. Pools are fixed
by the positions relative to the sequence's first one, so the layout now lives
on the memory and a ubatch only appends to it.

A sequence edit no longer stales every pooled key either, only the ones at or
after the edited position, which makes a tail seq_rm free. The pooling subgraph
is built unconditionally so the graph shape no longer changes every kpool
tokens, and the pool axis is folded into rows before soft_max, which otherwise
exceeds the CUDA gridDim.y limit past n_kv 262144.

Assisted-by: Claude Opus 5
(cherry picked from commit 5d1c40b93e17fddbf73b785efe43e0d02ccb3977)

* model : write the GLM5-Next recurrent rollback checkpoints

The conv state and the delta net state were only written to the live row, so a
rollback restored whatever the checkpoint rows happened to hold. Take the same
route as kimi-k3: build_recurrent_attn for the state, and write all K_rs conv
groups. That also drops a state view that assumed contiguous rows.

Enroll the arch in test-recurrent-state-rollback, which catches this under its
garbage-filled cache pass.

Assisted-by: Claude Opus 5
(cherry picked from commit 5ace37e86d5d448e83ef5dde5632c748185b18cd)

* llama: fix PR #27773 test-save-load-state restore failure

Clear the attention and indexer cache data after a failed hybrid state restore so restored NaNs cannot affect a later sequence.

Assisted-by: Codex

* llama: fix PR #27773 gpu-rocm graph reallocation

Reserve the full GLM5-Next pool capacity and dirty pool count. The gpu-rocm Test step aborts when n_new grows while the graph node count stays fixed; CUDA, Vulkan, Metal, and WebGPU checks report the same error.

Assisted-by: Codex

* llama : fix GLM5-Next k-pool layout staleness after edits and shared teardown

Two defects in the cross-ubatch k-pool layout added by the k-pool commit:

1. Wrong results. An edited sequence only rebuilt its pool layout when its cell
   count changed, so if the first ubatch after an edit added back exactly as many
   cells as were removed, the stale position-to-cell list survived. With a unified
   cache and more than one sequence, where another sequence takes the freed cells,
   the reused layout points at the wrong cells (CPU: large logit drift, CUDA: NaN).
   Rebuild whenever the sequence is stale, not only on a size mismatch.

2. Slowdown. "shared" mode was assumed to end only with an edit that forces a
   rebuild, but sharing also ends when the other sequence is removed. The survivor
   kept shared = true, pinning cache_safe off and re-pooling every pool on every
   ubatch (server trigger: n>1 completions with -kvu, via the seq_cp in
   copy_state_to). In seq_rm, if the layout has shared cells, stale every sequence
   so one rebuild re-derives sharing and cache_safe returns to 1.

Assisted-by: Claude Opus 5

* llama : fix build_attn_mha stream stride for non-contiguous q

build_attn_mha split the batch into streams with a stream stride of
q->nb[3]/n_stream. That only equals one stream's span, (ne[2]/n_stream)*nb[2],
when q is contiguous. GLM5-Next is nope-only, so it does not concat a rope part
and passes the permuted q_absorbed straight in, where nb[3] != ne[2]*nb[2]; the
stride was then n_head times too large and every stream s >= 1 read another
head's queries. Split-KV (-np N without --kv-unified) multi-stream prefill was
wrong for every stream past the first. Unified KV and decode were unaffected
(n_stream == 1, and decode takes the gather path). Other MLA models concat rope
so q is contiguous and the computed value is unchanged for them.

Compute the stride from the token dimension, which is identical for a
contiguous q.

Assisted-by: Claude Opus 5

* llama : re-derive GLM5-Next k-pool sharing on state_read/state_drop

The shared-cell teardown added to seq_rm (stale every sequence when the layout
has shared cells, so a survivor does not keep shared = true and pin cache_safe
off) was missing from the other paths that can free shared cells: state_read
and state_drop staled only the one sequence. Apply the same re-derivation there
and correct the comment that claimed sharing ends only via an edit or seq_rm.

Assisted-by: Claude Opus 5

* quant : drop duplicate GLM5-Next hc_ filter

The hc_ name filter was listed twice in the GLM5_NEXT protection block.

Assisted-by: Claude Opus 5

* glm5-next: scope K-pool cache access to indexed operations

* glm5-next: keep K-pool access in hybrid index memory

* glm5-next: keep mHC graph builders model-local

* glm5-next: mark only touched pools per ubatch

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Piotr Wilkin <ilintar@gmail.com>
2026-09-30 14:20:32 +08:00
Emanuil Rusev ba0ba54d93 server : remove the built-in UI's service worker when the UI is not served (#29565)
With --path or --no-ui, /sw.js returned 404, and a 404 does not remove a service worker, so browsers kept showing the cached built-in UI. Serve a worker that unregisters itself, clears its caches and reloads open tabs. A sw.js in the --path folder is still served first.

Assisted-by: Claude Opus 5.5
2026-09-29 17:48:17 +02:00
Adrien Gallouët 00af63567a common : use fs::path for config dir (#29649)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-29 17:22:49 +02:00
Georgi Gerganov c85b92c69c tests : adjust server string regex to also match m2 utlra results (#29648) 2026-09-29 15:32:35 +03:00
Georgi Gerganov 6d78fb0727 llama : fix init in several tools/examples (#29632) 2026-09-29 10:32:05 +03:00
Adrien Gallouët 76a5bc86d1 common : use fs::path for cache dirs (#29595)
- Avoid useless string conversions on Windows.
- No need for BSD or emscripten special cases.

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-29 07:24:03 +02:00
Georgi Gerganov 46e17a6352 tests : skip pytest workers when PYTEST_WORKERS=1 (#29610)
Assisted-by: pi:llama.cpp/Qwen3.8-27B
2026-09-29 08:20:54 +03:00
Pascal 526c43b8f7 mtmd: fix GCC 15 stringop-overflow in decode_embd_batch (#29607) 2026-09-29 01:21:37 +02:00
680a036285 server : support typed content (vision/audio/video) input for /v1/embeddings endpoint (#29556)
* server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding)

Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.

The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.

Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.

Assisted-by: Opencode Qwen3.8 27B

* clean up comments and docs

* refactor

* add tests

* support video and audio inp

---------

Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
2026-09-28 21:40:38 +02:00
Pascal 57b557cb95 models: pad on the left with ggml_pad_ext (#29567)
* models: pad on the left with ggml_pad_ext

The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders
build a left padding as a right pad followed by a roll, and DFlash2
concatenates a zero filled block in front of the previous tokens.
ggml_pad_ext does both in one node now that every backend supports a
left padding. The Gemma 4 audio embeddings are bit identical.

* models: skip the DFlash2 taps that only read padding

A tap at or past block_size shifts every row out of the block, so its
term is zero. The loop runs min(kernel_size, block_size) taps.
2026-09-28 20:56:15 +02:00
Xuan-Son Nguyen f1ea206218 batch: migrate speculative, mtmd and server to batch_ext (#29385)
* adapt common

* add common_batch

* wip

* wip: spec

* cont

* common_speculative_process

* server_batch to use common_batch

* rm some stale calls

Assisted-by: Claude Fable 5.1

* migrate mtmd

* handle imrope, handle return val of add()/add_embd()

* add spec zeros vector

* add warning on zero fill path
2026-09-28 19:52:45 +02:00
Sarah Wu 0c6a6a7ce5 Enables Windows ARM64 build with MSVC cl.exe (#28362)
* can reproduce the issue vlad sees

* fix fma issue

* drop volatile

* fix volatile runtime task

* add arm flag if needed

* fix hsum compile error

* fix syntax in quants

* strengthen sve probing

* make the syntax fixes one liners

* remove debug code

* formatting

* remove macro for float

* drive down gcc instruction count

* support armec

* fix CI comments

address CI comments

fix cross compile issue

remove warning

fix style and fix fma probing

fix style

* add documentation

* update documentation
2026-09-28 10:07:27 +02:00