The discover sidebar lists running downloads above the suggested models,
and steps aside while a query is active. A paused download stays in the
manager table, where it is resumed or dropped.
Assisted-by: pi
A stopped download keeps its partial file instead of deleting it, so
re-posting the tag resumes from disk. The first progress record now
carries the resume offset, the paused snapshot survives into the
resumed entry, and a pause settles as soon as it is requested.
Assisted-by: pi
The flag gates browsing and downloading HuggingFace GGUF models, so it is
declared on the discover layer; the repo-tree lookup in the manager layer
above reads it from here.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The manager dialog gains its second view: the Discover Models call to
action fades the table out, slides the title with its compass mark, and
fades the discover surface in; the arrow button returns to the table.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The Discover surface: the curated-catalog list with its search and
skeletons, the details pane with its readme, metadata, Hub stats and
download options, and the standalone download progress bar that replaces
the plain one. The quant download button sizes its select with a new xs
trigger, so the select gains that size and the muted box look. The store
stays incomplete when a repo fails, so the next mount retries it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
Add an allowHtml prop to MarkdownContent: raw HTML found in the markdown is
rendered after DOMPurify sanitization instead of being escaped as literal
text. Default stays escaped.
Assisted-by: pi:GLM-5.3-Flash
A Q4_0-mtp style tag now resolves the sidecar file when no model file matches it, so a solo draft or mmproj download actually pulls the file. Cached sidecar files list as their own entries so the state survives a restart, and removing such a tag deletes only the sidecar.
Assisted-by: pi:zai-org/GLM-5.3
The pane no longer says the server withholds data: it says the metadata
and the chat template appear once the model is loaded, and the template
block grows with its content instead of scrolling inside a fixed box. The
pane ends with a Delete this model from disk action, confirmed by the same
dialog the table row uses and hovering like a destructive dropdown entry.
The filter toggles list only the modalities the models at hand actually
carry, and starting a new chat from the pane closes the dialog.
Assisted-by: pi
The models table lists in-flight and paused downloads in a Downloading
section: flat rows like Loaded models, right after them, with a full
width progress bar along the bottom. The status column shows the percent
and swaps in pause or resume on hover; pause, resume and delete live in
the row's action menu. The selector lists running downloads only.
Assisted-by: pi
Turn the Hugging Face Hub models-metadata setting on by default, and when
it is off render no org avatar at all instead of the monogram fallback:
ModelOrgAvatar is the choke point, ModelAvatar stops mounting its wrapper,
and the base-model lookups in the selector and the download row are skipped.
Drop the vestigial "Enable Discover Models" setting from this branch: the
discover UI lives on the discover branch, and nothing here referenced it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The add menu keeps to what it attaches to the conversation, so the MCP
servers dialog moves to the sidebar rail and its menu entries and the
form callback they used go with it.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The reasoning level becomes its own control next to the selector, and
the add menu drops its reasoning submenu.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The selector keeps working when a provider is down: the trigger shows the
provider's mark and the org avatar, the banner only appears when no
backend is enabled, and the model ids read through their settings. The
model mark becomes MODEL_ICON everywhere.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
The selector surfaces move under models/ModelsSelector with their own
barrel: the dropdown and sheet relocate, the list, option and trigger icon
split out, and the shared list helpers join the navigation utils. The
searchable dropdown gains a sticky footer and per-surface class hooks.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
A manage models dialog for the local server: a table that folds each repo's
quants into one row, groups them into families, and sorts from its column
headers. Loaded models lead, then favorites, then the local block, then the
hidden one; sections keep their open state in local storage. The toolbar
filters by capability, modality and context, rows act on favorite, delete
and hide, and the manager opens focused on a model from the sidebar or a
download row. The models store gains the recents, hidden and group-open
state, and warms the Hub records for the context and capability columns.
Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash
* init conversion
* convert: ok
* model loaded
* add server code
* improve conversion script
* support shared prompt prefix
* add docs, imorove UX a bit
* add vision support
* add openjev tiny model for testing
* add dev docs
* support lev & kev
* clean up
* fix lev noul
* fix py lint
* nits docs
* clarify about not supporting date_facts
* Adding wide-load mmvq for Q8_0 and esimd dmmv for q8_0
Assisted-by: Codex
* remove guard for q8_0
* remove docs
* Simplify by committing to clean code without fallback
* Add feature flag as requested
Assisted-by: Claude Opus 5
---------
Co-authored-by: cwriter <cwriter@localhost>
* ggml : add `alloc_buffer_n` to buffer type interface
Add alloc_buffer_n method to ggml_backend_buffer_type_i
interface, with a public API ggml_backend_buft_alloc_buffer_n.
- Default implementation in ggml-backend.cpp handles multi-buffer
splitting and tensor allocation via ggml_tallocr
- Meta buffer type provides custom implementation that creates
per-device sub-contexts and delegates to simple buffer types
- ggml_backend_alloc_ctx_tensors_from_buft now collects tensors
into a list and delegates to the new API
- Remove temporary ggml_backend_meta_alloc_ctx_tensors_from_buft
- Add NULL alloc_buffer_n to all existing buffer type
interfaces (cpu, metal, openvino, hexagon, webgpu, zdnn, virtgpu, repack)
Assisted-by: llama.cpp:local pi
* cont : fix `cur_buf_size` init after flushing a buffer
* ggml : add TODO tag for shared buffer split logic
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
* tests : add alloc_buffer_n coverage
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
* cont : fix compile warnings
* tests : add descriptions for alloc_buffer_n tests
Assisted-by: pi:llama.cpp/Qwen3.8-27B
* ggml : address review comments on alloc_buffer_n
- restore GGML_LOG_ERROR on buffer alloc / tensor init failure in the
default impl (name the failing tensor)
- check the malloc result and drop the _impl indirection in
ggml_backend_alloc_ctx_tensors_from_buft
- remove comments that restate the code
- fix the TAG_ALLOC_SHARED_BUFFER_SPLIT typo
Assisted-by: pi:llama.cpp/Qwen3.8-27B
* ggml : add get_alloc_size_n to buffer type interface
- Add ggml_backend_buft_get_alloc_size_n public API
- Add optional get_alloc_size_n callback to ggml_backend_buffer_type_i
- Share tensor->buffer planning between alloc_buffer_n default and get_alloc_size_n default
- Replace unchecked realloc with std::vector in alloc_buffer_n default
- Make ggml_backend_alloc_ctx_tensors_from_buft_size use the new API
- Add test-alloc coverage for get_alloc_size_n
Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-Vision-Exp
* cont : report malloc failure
In tool.uv.sources, torch was unconditionally pinned to the custom
pytorch CPU index, which lacks macOS Darwin wheels and causes uv sync
to fail on macOS. Add the sys_platform == 'linux' marker to match the
existing Poetry dependencies configuration.
Assisted-by: Antigravity
Resolves: https://github.com/ggml-org/llama.cpp/issues/29176
* hexagon: add q2_k and q3_k quant type support
* hex-qk: consistent allocation of src1_row_size
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* tests : simplify function signature
* llama : clamp kpool re-pool bound to existing pools
The n_tokens/kpool + n_seqs_unq bound on n_new_g overshoots when a batch
fills the whole cache: n_ctx tokens complete exactly n_ctx/kpool pools, so
the +1 pads new_pool_idxs/new_pool_rep one entry past n_pool_real. Graph
reserve only covers n_pool_real entries, so the first full-context decode
builds bigger tensors than reserved and ggml-alloc demands a graph
reallocation (abort under GGML_SCHED_DEBUG_REALLOC=1).
Clamp the bound to n_pool_real: a ubatch can never mark more pools than
the cache holds, and reserve's n_pool_max already covers that.
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-MOPD
* cont : cap to n_pool_max
* hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA
* hex-cpy: various fixes on top of the concat optimizations
Removed CONCAT_DMA_MIN_ROW logic, it was broken with 64-bit DMA.
While it's kinda silly to use DMA for tiny stuff if that tensor gets mapped to an extended buffer the only way to read it is DMA.
Added missing dma_queue_flush() calls.
Added additional guards for conditions we don't support.
---------
Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>
* convert : write Gemma embedding scale for DFlash drafts
A DFlash draft shares the target's token embeddings. Gemma scales them by sqrt(hidden_size) in the forward pass, and the draft config does not state that scale, so the converted draft read unscaled embeddings.
Take the scale from the target config when the draft config has none.
Assisted-by: Claude
* convert : check with get_model_architecture for gemma models
* cuda : route sm70 to the Turing MMVQ nwarps table
Volta (sm_70) has no MMVQ parameter table of its own and falls through
to GENERIC, which launches K-quant batch-1 decode (ncols_dst == 1) at
nwarps=4. sm_70 shares TURING's tuning: the K-quant vec_dot prefers
nwarps=2 there. Route sm_70 to the existing MMVQ_PARAMETERS_TURING
table in both the device and the host table selector.
Measured on one Tesla V100 32GB PCIe (PG500-216, driver 580.178.04,
CUDA 12.0.140) with Qwen3.8-27B Q4_K_M, tg128, interleaved A/B in 6
ABBA blocks with paired per-block deltas: +1.091 t/s = +3.17 %
(t = +49.0, all six per-block deltas positive); perplexity
bit-identical (6.3697 +/- 0.04066 both builds, wiki.test.raw). The
patched build's K-quant mul_mat_vec_q kernels launch at nwarps=2
(cubin EIATTR_MAX_THREADS) while Q4_0/Q8_0 stay at nwarps=4, and the
same measurement on the September master base gave +3.84 % (t = 85).
The tuning originates from the V100-focused fork anyei/llamacpp-v100
(MIT), commit b912d1b1e, which carries a dedicated
MMVQ_PARAMETERS_VOLTA table; a cubin-level comparison confirmed that
routing sm_70 to the existing TURING table is equivalent for the
K-quant batch-1 path this change affects, so this is the minimal
2-line form. https://github.com/anyei/llamacpp-v100/commit/b912d1b1e
Original-patch-by: anyei <angelyoelroblesmercedes@gmail.com>
* Update ggml/src/ggml-cuda/mmvq.cu
---------
Co-authored-by: tkittich <tkittich@gmail.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* metal : release temporary private transfer buffers
Assisted-by: OpenAI Codex
* metal : fix order and formatting
---------
Co-authored-by: Niklas Wenzel <dev@nikwen.de>
* CUDA: Handle compute type for NVFP4 on cublass path
Signed-off-by: ynankani <ynankani@nvidia.com>
* Use BF16 compute type for quantized models if HW allows
Signed-off-by: ynankani <ynankani@nvidia.com>
* Set acc prec to bf16 for nvfp4 as it needs atleast bf16 range
Signed-off-by: ynankani <ynankani@nvidia.com>
* Update ggml/src/ggml-cuda/ggml-cuda.cu
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* preserve op_params for per-expert matmul
Signed-off-by: ynankani <ynankani@nvidia.com>
---------
Signed-off-by: ynankani <ynankani@nvidia.com>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Pinning to >= 3.4.3 is required to enable DeviceTopK, which was affected by
a race condition https://github.com/NVIDIA/cccl/pull/10627.
We will relax this for future CTK versions which will bundle CCCL >
3.4.X (CTK 13.5 will bundle CCCL 3.5.0 for example)