The add menu keeps to what it attaches to the conversation, so the MCP
servers dialog moves to the sidebar rail and its menu entries and the
form callback they used go with it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
Adds a measured-px height transition to the tool category groups in the
Tools submenu, replacing the bits-ui Collapsible whose conditional
rendering kills the transition. The same technique as the reasoning panel:
the rows stay mounted, the height animates between 0 and the measured
scrollHeight, and visibility keeps collapsed rows out of the tab order.
Assisted-by: pi:GLM-5.3-Flash
Serve the recommended MCP server favicons and the Hugging Face badge from
the app's base path, not the domain root, so they resolve when the app is
mounted under a path. Also normalizes the safe HTML config indentation.
Assisted-by: pi:GLM-5.3-Flash
A row click slides the pane in as a fixed-width drawer: the toolbar's
calls to action leave first, the pane is laid out before its first open,
and a model switch fades one out and the next in. The model information
dialog folds into the pane's information tab, and a chat started from
the pane closes the manager and focuses the composer.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The pane shows a model's information, load and inference settings side by
side: what its provider reports, the Hub records it falls back to, the
load controls gated by what the serving backend can do, and the
per-model overrides the load form writes. The slider control arrives
with it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The table grows a section per enabled provider, each with its own mark,
error and loading state, and a provider filter with live repo counts.
Rows gain their capability gates back, the draft column with its
use-as-draft action, and the quant badge names the provider on its rows.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
A provider is added from a preset card or by hand: the dialog probes the
endpoint, reads a refusal as a llama.cpp server that wants a key, and
shows the connection test as a status block. Presets carry their official
artwork and a saved backend keeps its branding through the favicon
fallback. The manager dialog gains its third view, and the settings save
stops clobbering keys that other surfaces write.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
Chat goes through protocol adapters, so an OpenAI-compatible endpoint
speaks its own wire format: per-backend paths and headers, the model on
the request, tools kept on the local server, and token counts synthesized
for endpoints that do not stream their own timings. The server store keeps
the local server's props while another provider is active, a conversation
resolves the provider its model belongs to before sending, and the chat
screen never blocks on the local probe when the install has none.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
A provider is a server entry with a base url, an optional key, a protocol
and the paths its API lives at; the local llama.cpp server is the built-in
one. Backends persist in settings and the active one is restored on load.
Requests resolve against the active backend, model ids become
backend-qualified, and every backend's model list is fetched and cached in
the background. The manager's helpers learn to read a model's drafts,
context and the provider that serves it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The flag gates browsing and downloading HuggingFace GGUF models, so it is
declared on the discover layer; the repo-tree lookup in the manager layer
above reads it from here.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The manager dialog gains its second view: the Discover Models call to
action fades the table out, slides the title with its compass mark, and
fades the discover surface in; the arrow button returns to the table.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The Discover surface: the curated-catalog list with its search and
skeletons, the details pane with its readme, metadata, Hub stats and
download options, and the standalone download progress bar that replaces
the plain one. The quant download button sizes its select with a new xs
trigger, so the select gains that size and the muted box look. The store
stays incomplete when a repo fails, so the next mount retries it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.
Assisted-by: pi:GLM-5.3-Flash
A Q4_0-mtp style tag now resolves the sidecar file when no model file matches it, so a solo draft or mmproj download actually pulls the file. Cached sidecar files list as their own entries so the state survives a restart, and removing such a tag deletes only the sidecar.
Assisted-by: pi:zai-org/GLM-5.3
Turn the Hugging Face Hub models-metadata setting on by default, and when
it is off render no org avatar at all instead of the monogram fallback:
ModelOrgAvatar is the choke point, ModelAvatar stops mounting its wrapper,
and the base-model lookups in the selector and the download row are skipped.
Drop the vestigial "Enable Discover Models" setting from this branch: the
discover UI lives on the discover branch, and nothing here referenced it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The add menu keeps to what it attaches to the conversation, so the MCP
servers dialog moves to the sidebar rail and its menu entries and the
form callback they used go with it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The reasoning level becomes its own control next to the selector, and
the add menu drops its reasoning submenu.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The selector keeps working when a provider is down: the trigger shows the
provider's mark and the org avatar, the banner only appears when no
backend is enabled, and the model ids read through their settings. The
model mark becomes MODEL_ICON everywhere.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The selector surfaces move under models/ModelsSelector with their own
barrel: the dropdown and sheet relocate, the list, option and trigger icon
split out, and the shared list helpers join the navigation utils. The
searchable dropdown gains a sticky footer and per-surface class hooks.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
A manage models dialog for the local server: a table that folds each repo's
quants into one row, groups them into families, and sorts from its column
headers. Loaded models lead, then favorites, then the local block, then the
hidden one; sections keep their open state in local storage. The toolbar
filters by capability, modality and context, rows act on favorite, delete
and hide, and the manager opens focused on a model from the sidebar or a
download row. The models store gains the recents, hidden and group-open
state, and warms the Hub records for the context and capability columns.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* convert : write Gemma embedding scale for DFlash drafts
A DFlash draft shares the target's token embeddings. Gemma scales them by sqrt(hidden_size) in the forward pass, and the draft config does not state that scale, so the converted draft read unscaled embeddings.
Take the scale from the target config when the draft config has none.
Assisted-by: Claude
* convert : check with get_model_architecture for gemma models
* cuda : route sm70 to the Turing MMVQ nwarps table
Volta (sm_70) has no MMVQ parameter table of its own and falls through
to GENERIC, which launches K-quant batch-1 decode (ncols_dst == 1) at
nwarps=4. sm_70 shares TURING's tuning: the K-quant vec_dot prefers
nwarps=2 there. Route sm_70 to the existing MMVQ_PARAMETERS_TURING
table in both the device and the host table selector.
Measured on one Tesla V100 32GB PCIe (PG500-216, driver 580.178.04,
CUDA 12.0.140) with Qwen3.8-27B Q4_K_M, tg128, interleaved A/B in 6
ABBA blocks with paired per-block deltas: +1.091 t/s = +3.17 %
(t = +49.0, all six per-block deltas positive); perplexity
bit-identical (6.3697 +/- 0.04066 both builds, wiki.test.raw). The
patched build's K-quant mul_mat_vec_q kernels launch at nwarps=2
(cubin EIATTR_MAX_THREADS) while Q4_0/Q8_0 stay at nwarps=4, and the
same measurement on the September master base gave +3.84 % (t = 85).
The tuning originates from the V100-focused fork anyei/llamacpp-v100
(MIT), commit b912d1b1e, which carries a dedicated
MMVQ_PARAMETERS_VOLTA table; a cubin-level comparison confirmed that
routing sm_70 to the existing TURING table is equivalent for the
K-quant batch-1 path this change affects, so this is the minimal
2-line form. https://github.com/anyei/llamacpp-v100/commit/b912d1b1e
Original-patch-by: anyei <angelyoelroblesmercedes@gmail.com>
* Update ggml/src/ggml-cuda/mmvq.cu
---------
Co-authored-by: tkittich <tkittich@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* metal : release temporary private transfer buffers
Assisted-by: OpenAI Codex
* metal : fix order and formatting
---------
Co-authored-by: Niklas Wenzel <dev@nikwen.de>
* CUDA: Handle compute type for NVFP4 on cublass path
Signed-off-by: ynankani <ynankani@nvidia.com>
* Use BF16 compute type for quantized models if HW allows
Signed-off-by: ynankani <ynankani@nvidia.com>
* Set acc prec to bf16 for nvfp4 as it needs atleast bf16 range
Signed-off-by: ynankani <ynankani@nvidia.com>
* Update ggml/src/ggml-cuda/ggml-cuda.cu
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* preserve op_params for per-expert matmul
Signed-off-by: ynankani <ynankani@nvidia.com>
---------
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.
We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)
* BLAS : Document AOCL-BLAS build and label the device AOCL-BLAS
* AOCL-Blas : Add an AOCL-BLAS Quick Start and drop the fixed version path
* AOCL-BLAS doc : Note on ZenDNN
LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space
after every special token and parallel tool calls are separated by
<|end|>. The GPT-OSS handler rejects this output, so add a dedicated
handler, selected by the chat_format=llm-jp-harmony-v1 declaration in
the chat template.
Assisted-by: Claude Fable 5.1
* vocab : honor BOS/EOS settings for PLaMo-2 and PLaMo-3
The original tokenizer configs for PLaMo-2 and PLaMo-3 have
`add_bos_token: true` and `add_eos_token: false` , but
_set_vocab_plamo() did not write the BOS/EOS metadata. The
PLAMO2 tokenizer path also ignored add_bos/add_eos during
tokenization.
Write the settings from tokenizer_config.json and honor them in
the PLAMO2 tokenization path. GGUFs without these keys keep the
previous behavior.
* Update conversion/base.py
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
The committed docs/ops/CPU.csv is out of sync with the current
test-backend-ops suite: 11 ops with CPU support (COL2IM_1D,
MUL_MAT_HADAMARD, SWIGLU_CLAMP, MUL_MAT_W4A4/W4A8, MUL_MAT_ID_W4A4/W4A8,
DSV4_HC_COMB/PRE/POST, LIGHTNING_INDEXER) are missing entirely, and
many other ops have fewer test cases than the suite generates now.
docs/ops.md (which CI requires to match the CSVs) therefore
understates CPU support.
Regenerated with:
test-backend-ops support -b CPU --output csv > docs/ops/CPU.csv
scripts/create_ops_docs.py
Note: ADD1 now reads unsupported on CPU because ggml_add1 is
GGML_DEPRECATED and the suite no longer generates test cases for it;
the CPU implementation itself is still present.
Assisted-by: Xing
#27941 disabled -sm tensor for qwen4exp because test-llama-archs asserted on the
Meta device once the fixture carried a PLE layer:
GGML_ASSERT(ggml_backend_buffer_is_meta(tensor->buffer)) at ggml-backend-meta.cpp:476.
With host-resident embeddings the PLE gather is a CPU node and hc_init (the REPEAT
that fans the embedding out to the hc streams) was first reached through layer 0's
PLE path, after that gather. ggml_backend_sched_split_graph pass 2 expands a device
assignment upwards only until it meets a CPU node, so the REPEAT stayed on the CPU
and the later reshape of hc_init inside the meta split viewed a host-resident node.
Expanding hc_init right after it is built puts the REPEAT directly before the first
device node, where pass 2 assigns it; the embedding reshape stays in the CPU split
and is copied in as a split input, as in deepseek4.
dequantize_block_iq4_nl writes QK_K values per block, but a row can be shorter than that (an IQ4_NL row is only guaranteed to be a multiple of QK4_NL). Threads whose 32-value sub-block starts at or past k currently read and write past the end of the row. Skip those sub-blocks; for rows that are a multiple of QK_K the check never fires.
* hex-allreduce: add support for safe scatter mode
* hex-allreduce: pare down excessive comments
* hex-allreduce: re-write to remove register spills
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>