Redirects for registry and blob transfers now validate the target scheme
and resolved addresses before following, re-check DNS on each redirect,
and do not follow redirects that switch an https session to plain http.
The --insecure option continues to relax address checks for private
registries but not scheme checks.
This is a rewrite of the create functionality for the MLX engine.
The core idea behind the create functionality is to break the import/convert into a pipeline of distinct phases:
* Read (scan the safetensors directory for the various bits of metadata)
* Classify (determine what the import type)
* Plan (determine any transforms that need to be done)
* Write (transform any data as necessary and write out the blobs)
* Create the manifest
Each architecture has a "policy" which determines how to convert the model correctly. A number of different formats for safetensors are supported including:
* nvfp4 (two formats: model optimized, torch)
* fp8 datatypes (convert to mxfp8)
* standard bf16 based weights
A number of cleanups/simplifications have been done including:
* using the baked in names for the tensors instead of munging them into something else
* unified 3d expert tensors (instead of separate per expert tensors)
* fewer unnecessary transforms to the various tensors in a model (keep a model as close to the source as possible)
* unified capability checking
* draft model handling (for MTP) is done on the same path
Image generation has been intentionally removed.
This change allows .experts.gate_proj / .up_proj / .down_proj tensor names to each
be used for both quantized (i.e. nvfp4 and mxfp8) and non-quantized (bf16) models.
Previous to this only non-quantized models used that tensor naming scheme.
Adding/Multiplying a tensor by a scalar w/ a different data type
can cause the tensor to be promoted and cause performance issues.
This change adds several guards against over-promotion.
This change addresses some problems with GGUF conversion including:
* correctly naming the MoE tensors
* correctly quantizing the nextn.eh_proj.weight MTP tensor
This change updates the show API for MLX models to:
* display the correct quantization in mixed precision models
* not display the global_scale scalar value
* not duplicate the `tools` capability
This change adds dflash block diffusion speculative decoding to the MLX runner. Included in this change:
support for qwen3.6 moe/dense speculative decoding
draft model recurrent cache playback
RoPE/YaRN changes (DRY out the laguna/dflash MoE YaRN implementation)
support for greedy sampling / leviathan/chen sampling
* mlx: rework the MLX sampler
Replace the MLX sampler transform chain with an explicit distribution pipeline that applies:
1. penalties
2. top-k
3. temperature/softmax
4. top-p
5. min-p
6. normalize
7. categorical
The common top_k path now keeps sparse [B,K] token ids/probabilities on GPU instead of carrying full-vocab
scores, and sampled MTP reuses those draft/target distributions for acceptance, bonus, and residual sampling.
This change also fixes the seed parameter so that temperature sampling and sampled MTP are reproducible.
This change adds support for MTP (multi-token prediction) speculative decoding for the
gemma4 model family.
It includes:
* support for importing safetensors based gemma4 draft models with `ollama create`
* a new DRAFT command in the Modelfile for specifying draft models
* a --quantize-draft flag for the ollama create command to quantize the draft model
* cache support for speculation
* changes to the rotating cache to be able to handle MTP correctly
* sampling support for draft model token prediction
---------
Co-authored-by: Daniel Hiltgen <daniel@ollama.com>
This change fixes two issues with Modelfiles:
1. If a user uses `ollama show --modelfile` to show a safetensors based
model, the Model would leave the "FROM" field blank which won't allow
a user to recreate the model. This change adds the model's current
canonical short name to the FROM field.
2. If a user uses the `/save` command in the CLI any messages which were
saved in a previous model wouldn't get saved (only the set of messages
from the current session).