The secondary button gets its own look back and the chat add button its
own light surface. Pressable elements get a pointer cursor again, the
font rendering smooths on the app shell, the root layout resolves its
props probe through the providers' api url, and the agent skills stay
out of prettier's way.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The add menu keeps to what it attaches to the conversation, so the MCP
servers dialog moves to the sidebar rail and its menu entries and the
form callback they used go with it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
A row click slides the pane in as a fixed-width drawer: the toolbar's
calls to action leave first, the pane is laid out before its first open,
and a model switch fades one out and the next in. The model information
dialog folds into the pane's information tab, and a chat started from
the pane closes the manager and focuses the composer.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The pane shows a model's information, load and inference settings side by
side: what its provider reports, the Hub records it falls back to, the
load controls gated by what the serving backend can do, and the
per-model overrides the load form writes. The slider control arrives
with it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
Adds a measured-px height transition to the tool category groups in the
Tools submenu, replacing the bits-ui Collapsible whose conditional
rendering kills the transition. The same technique as the reasoning panel:
the rows stay mounted, the height animates between 0 and the measured
scrollHeight, and visibility keeps collapsed rows out of the tab order.
Assisted-by: pi:GLM-5.3-Flash
Serve the recommended MCP server favicons and the Hugging Face badge from
the app's base path, not the domain root, so they resolve when the app is
mounted under a path. Also normalizes the safe HTML config indentation.
Assisted-by: pi:GLM-5.3-Flash
The reasoning level becomes its own control next to the selector, and
the add menu drops its reasoning submenu.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The selector keeps working when a provider is down: the trigger shows the
provider's mark and the org avatar, the banner only appears when no
backend is enabled, and the model ids read through their settings. The
model mark becomes MODEL_ICON everywhere.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The selector surfaces move under models/ModelsSelector with their own
barrel: the dropdown and sheet relocate, the list, option and trigger icon
split out, and the shared list helpers join the navigation utils. The
searchable dropdown gains a sticky footer and per-surface class hooks.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The table grows a section per enabled provider, each with its own mark,
error and loading state, and a provider filter with live repo counts.
Rows gain their capability gates back, the draft column with its
use-as-draft action, and the quant badge names the provider on its rows.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
A provider is added from a preset card or by hand: the dialog probes the
endpoint, reads a refusal as a llama.cpp server that wants a key, and
shows the connection test as a status block. Presets carry their official
artwork and a saved backend keeps its branding through the favicon
fallback. The manager dialog gains its third view, and the settings save
stops clobbering keys that other surfaces write.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
Chat goes through protocol adapters, so an OpenAI-compatible endpoint
speaks its own wire format: per-backend paths and headers, the model on
the request, tools kept on the local server, and token counts synthesized
for endpoints that do not stream their own timings. The server store keeps
the local server's props while another provider is active, a conversation
resolves the provider its model belongs to before sending, and the chat
screen never blocks on the local probe when the install has none.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
A provider is a server entry with a base url, an optional key, a protocol
and the paths its API lives at; the local llama.cpp server is the built-in
one. Backends persist in settings and the active one is restored on load.
Requests resolve against the active backend, model ids become
backend-qualified, and every backend's model list is fetched and cached in
the background. The manager's helpers learn to read a model's drafts,
context and the provider that serves it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The manager dialog gains its second view: the Discover Models call to
action fades the table out, slides the title with its compass mark, and
fades the discover surface in; the arrow button returns to the table.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The Discover surface: the curated-catalog list with its search and
skeletons, the details pane with its readme, metadata, Hub stats and
download options, and the standalone download progress bar that replaces
the plain one. The quant download button sizes its select with a new xs
trigger, so the select gains that size and the muted box look. The store
stays incomplete when a repo fails, so the next mount retries it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.
Assisted-by: pi:GLM-5.3-Flash
A Q4_0-mtp style tag now resolves the sidecar file when no model file matches it, so a solo draft or mmproj download actually pulls the file. Cached sidecar files list as their own entries so the state survives a restart, and removing such a tag deletes only the sidecar.
Assisted-by: pi:zai-org/GLM-5.3
A manage models dialog for the local server: a table that folds each repo's
quants into one row, groups them into families, and sorts from its column
headers. Loaded models lead, then favorites, then the local block, then the
hidden one; sections keep their open state in local storage. The toolbar
filters by capability, modality and context, rows act on favorite, delete
and hide, and the manager opens focused on a model from the sidebar or a
download row. The models store gains the recents, hidden and group-open
state, and warms the Hub records for the context and capability columns.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The manager and later surfaces share one set of row parts: the collapsible
section and grouped list containers, the avatar with its org and quant
badges, the capability and context columns with their Hub fallbacks, the
load control, the shared row action set, and the download row with its
progress bar. The toggle and toggle-group controls and the tertiary button
variant arrive with them.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
hasBadges now counts draft sidecar badges, so a sidecar-only model still
renders. Billions keep one decimal for hub counts and stay bare for whole
values. Avatar failures track the org instead of the instance, and the
download progress bar no longer pulses while determinate.
Assisted-by: pi:zai-org/GLM-5.3-Flash
Extract ModelCapabilityIcons (canonical Tools/Reasoning/Vision/Video/Audio
order) out of ModelId and reuse it there, add the shared DialogConfirmDownload
for destructive download actions, the discover org avatar with dark-mode
inversion and the thin download progress bar, and rework ModelId badges to take
thinking/tool-use support directly.
Assisted-by: pi:GLM-5.3-Flash
Track HuggingFace downloads end to end: the server download/cancel endpoints,
a status manager fed by the /models/sse download progress events, and a
models-discover store holding the catalog and detail state for the discover
view. Downloaded and in-flight entries are excluded from the loadable model
list.
Assisted-by: pi:GLM-5.3-Flash
Replace the raw runtime-memory estimate with the app's compatibility check:
the smallest real Mac memory tier that fits a model file, budgeted as
RAM x 0.75 minus fixed overhead with headroom on the file size. The constants
move to lib; the unused runtime-memory estimate is dropped. browser-info's
OS detection is exported for reuse.
Assisted-by: pi:GLM-5.3-Flash
Carries the HTTP status on retryable fetch errors instead of matching the
message text. Marks expand-dependent catalog fields optional and documents
the data/models index pairing. Adds table tests for the pure helpers.
Assisted-by: pi:zai-org/GLM-5.3-Flash
Per review: drop JSDoc that restates the method name and inline comments
that restate the code; keep only comments carrying non-obvious context.
Assisted-by: pi:zai-org/GLM-5.3-Flash
Add HuggingFaceService and its constants/enums/types: GGUF repo search, file
tree and model detail fetching, quant/sidecar filename analysis, shard-set
collapsing and the llama.app catalog feed, plus an orgOf() helper on the model
name utils.
Assisted-by: pi:GLM-5.3-Flash
Extend the shared model id parser with sidecar tokens (draft variants and
auxiliary imatrix/mmproj files), weight-file and custom-quant regexes, and add
the tools capability to ModelCapabilities; the selector option row picks it up
from the model's declared capabilities.
Assisted-by: pi:GLM-5.3-Flash
* ci : update the oneAPI toolkit to 2026.1
oneDNN is removed from Intel Deep Learning Essentials in 2026.0, so
staying on the deep-learning-essentials path would silently lose oneDNN
support when the toolkit version is updated. Switch both the Ubuntu and
Windows CI jobs to the new unified Intel oneAPI Toolkit installer,
which still includes oneDNN (until 2027.0) and keeps the component IDs
unchanged for the Windows install script.
Measured with the same code (b10899) built with oneAPI 2026.1 vs the
2025.3-based release build on Arc B570: prompt processing 1331 vs 434
t/s (3.1x), token generation 50.1 vs 45.3-48.0 t/s.
Assisted-by: GLM (z-ai/glm-5.3-flash)
* docs : update the SYCL backend build requirements for oneAPI 2026.1
With the 2026.0 release the Base toolkit and the HPC toolkit are
combined into the oneAPI Toolkit, and oneDNN is removed from the Deep
Learning Essentials package. Update the install instructions, the
verified release table and the news section accordingly.
Assisted-by: GLM (z-ai/glm-5.3-flash)
* ci : update the release workflow for oneAPI 2026.1 and Level Zero SDK 1.33.1
Align the release package build with the CI build update:
- oneAPI toolkit 2025.3.3 -> 2026.1 (the unified oneAPI Toolkit)
- Level Zero SDK 1.28.2 -> 1.33.1, and the Debian package names
(level-zero/level-zero-devel -> libze1/libze-dev)
- The Windows DLL copy list for the 2026.1 runtime: sycl9.dll and the
.6/.3 MKL library versions
Assisted-by: GLM (z-ai/glm-5.3-flash)
* ci : remove the removed .spv fallback files from the Windows DLL copy list
oneAPI 2026.1 no longer ships libsycl-fallback-bfloat16.spv and
libsycl-native-bfloat16.spv (the OpenCL fallback mechanism changed), so
the copy step failed with exit 1.
Assisted-by: GLM (z-ai/glm-5.3-flash)
* devops : update the oneAPI toolkit image in the Intel Dockerfile
Assisted-by: GLM (z-ai/glm-5.3-flash)
---------
Co-authored-by: Asahi-Prv <Asahi-Prv@users.noreply.github.com>
* server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding)
Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.
The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.
Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.
Assisted-by: Opencode Qwen3.8 27B
* clean up comments and docs
* refactor
* add tests
* support video and audio inp
---------
Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
* models: pad on the left with ggml_pad_ext
The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders
build a left padding as a right pad followed by a roll, and DFlash2
concatenates a zero filled block in front of the previous tokens.
ggml_pad_ext does both in one node now that every backend supports a
left padding. The Gemma 4 audio embeddings are bit identical.
* models: skip the DFlash2 taps that only read padding
A tap at or past block_size shifts every row out of the block, so its
term is zero. The loop runs min(kernel_size, block_size) taps.
* adapt common
* add common_batch
* wip
* wip: spec
* cont
* common_speculative_process
* server_batch to use common_batch
* rm some stale calls
Assisted-by: Claude Fable 5.1
* migrate mtmd
* handle imrope, handle return val of add()/add_embd()
* add spec zeros vector
* add warning on zero fill path