Model load code eagerly evaluated every weight fold (expert stacking,
gather transposes, gate/up fusing) as it was built, with the folds
running on the GPU against lazily loaded tensors: Metal committed
command buffers that waited on file reads, and macOS kills command
buffers that stall too long, so loading a large model from a slow
volume aborted with "Command buffer execution failed". The eager evals
also kept every layer's fold sources alive until the post-load sweep,
transiently holding roughly twice the expert weights on MoE models.
Build the folds lazily and let the runner's weight eval run them, and
on Metal materialize the loaded tensors with CPU reads before any
weight graph exists: no command buffer is ever committed waiting on
file data, at any storage speed, and fold sources free as their folds
execute. CUDA loads read at dispatch and skip the pre-pass. Models no
longer evaluate weights at load; on Metal, tensors the model does not
retain are now read before the sweep frees them.
Measured on an M5 Max, warm page cache, greedy outputs bit-identical:
before after
nemotron-3.5-lightning:30b-mlx 1.9s 39.7GiB 1.45s 24.7GiB
qwen3.6:35b-mlx 1.27s 22.5GiB 1.1-1.2s 22.4GiB
nemotron, reads at ~60MB/s aborts in 6s loads in 346s
Fixes#17902
The MLX runner accepted the API's format field but did not enforce it:
requests asking for JSON or a JSON Schema got unconstrained text, and
clients had no way to tell.
Enforce format with xgrammar: each sampling step masks the logits to
the tokens the grammar allows, so every emitted token and the end of
generation are valid under the constraint. Sampling, penalties, and
logprobs see the constrained distribution, and "json" yields a JSON
object, as the API documents and the llama-server path already
enforces. Only sampling waits on the mask; the forward pass is
dispatched before it, so constrained decoding stays pipelined.
The grammar engine is a dynamic library alongside MLX; when it is
missing, plain inference is unaffected and structured requests fail
with an explicit error. Constrained requests decode without
speculative decoding for now.
Decoding 256 tokens of a book-list schema on qwen3.8:27b-mlx (M5 Max,
seed 42, thinking off); pre-decode is the request time spent before
the first token:
unconstrained ~65 tok/s pre-decode ~70 ms
unconstrained, no draft ~32 tok/s pre-decode ~70 ms
JSON schema ~32 tok/s pre-decode ~70 ms
Schema and draft-less decoding are equal to within 0.1 tok/s in
paired adjacent requests, and a cold grammar compile adds nothing
measurable to pre-decode. The gap to unconstrained decoding is the
disabled draft model.
Fixes#16563
Co-authored-by: Daniel Hiltgen <daniel@ollama.com>
Token ids are int32 throughout the runner, so every caller reading ids
out of an int32 array narrowed the widened value right back. Make Int
and Ints return int32 and Float return float32, matching Floats, and
require the exact dtype instead of accepting and widening every
integer and float width: no caller read anything through those paths
but int32 tokens.
Ints and Floats also copied out of the array's buffer without
evaluating it first, so reading an array still in flight after an
async dispatch could return unwritten data, and correctness depended
on every call site remembering an explicit Eval. Evaluate in every
reader, matching the scalar readers, which already wait through item.
An available array costs a status check and an in-flight one waits
for its event; only a never-dispatched array evaluates a graph.
Nothing has set Grammar since the CGO engine removal took its writers
out; it survived as a read-only pass-through on the llama-server path
and a comment claiming it is set before dispatch. Remove the field and
the dead pass-through. llama-server keeps its wire-level grammar field,
which the "json" format conversion still uses.
Claude Code adds a "tokens left" system message after every tool
result. Since ollama moves system messages to the front of the prompt,
this breaks the KV cache on every request.
A long prompt records restore points during prefill, but they only
reached the prefix trie when the prefill completed; a cancelled request
closed and released everything it had captured. Agent clients routinely
cancel long prefills — their timeouts are shorter than the minutes a
40k-token prompt takes — so every retry started the whole prompt over
and never got further than the timeout allowed, which presents as the
model hanging forever.
Closing a session now attaches every snapshot the prefill crossed, so a
retry resumes from the last one and makes progress across timeouts.
Scenario tests cover retries resuming exactly where a cancelled attempt
stopped and cancellations on divergent conversation variants.
Fixes#17839
A prefill that resumes partway into cached history — routine once
client timeouts interrupt long prompts — used to attach its captures
onto a node extended in place, so the stored snapshot spanned only the
tokens the prefill evaluated while the node's edge reached further
back. Restores walk node by node and trust each snapshot to cover its
node's edge; the short snapshot stranded the caches at mismatched
offsets and, on models with recurrent layers, ended up freeing all
cache state — a request matching 46k of a 47k-token prompt reprocessed
from zero.
Growth now never extends a node underneath its snapshots. New tokens
become a child node that carries exactly its own captures, and the
path stays compressed because non-user segments merge back into their
parent through the caches' snapshot Merge. Close already pages out
what it records, so every merge combines adjacent covered snapshots
and every stored snapshot spans exactly its node's edge.
When a session closes, every cache rests exactly at the end of the
segment the trie is about to record. That is the one moment the
segment's state can be captured for every layer, so close now pages
the new segment out itself instead of recording it without snapshots
and leaving the capture to a later path switch.
Path switching then has nothing left to capture and only rewinds and
pages in. The whole-state entry taken at close is released when the
next request grows past the segment; sliding-window layers pay the
same window copy a scheduled capture already costs.
Page-in restores a path node by node and trusts each stored snapshot to
cover its node's whole edge. A capture taken during prefill spans from
the previous capture or the prefill base, which need not line up with
the node it lands on: when a prefill resumes partway into cached
history, a capture can reach back before its node's start, and a
capture landing on a node that already has snapshots replaced them
with a shorter span that page-in then could not serve.
Clip each capture to its node's edge on attach, and keep the snapshots
the node already has instead of replacing them.
A prefill settles the drafter with the seed token after its last chunk,
leveling the draft caches with the targets; a cancelled prefill
returned before that, leaving the targets one token past the draft
caches and the recorded keys. The next request then had to move every
cache, and models with recurrent layers, which cannot rewind, fell back
to the last snapshot: a retry after a client timeout lost up to a full
snapshot interval of the prompt it had just evaluated.
Settle with the next prompt token on the cancelled path too. The caches
then rest level with the recorded keys, and a retry resumes exactly
where the prefill stopped.
When a builtin parser rejects model output, the completion callback wrote the
error to an unbuffered channel and returned. The callback cannot stop
generation -- it has no error return -- so the next chunk re-entered the
callback, hit the same parse error and blocked writing to a channel the
consumer had already stopped reading after emitting its 500. The completion
never returned, the goroutine leaked and the runner request was never
released, so retrying the same prompt hung with no log output until the client
gave up.
Record the parse error, cancel the completion, and report it once the
completion has returned. Parse failures landing on the final chunk were
already terminal, which is why non-thinking requests and the direct
qwen3-coder parser path failed cleanly and only thinking mode wedged.
ChatHandler and GenerateHandler share the defect: both run the same parser in
the same shape of callback behind a consumer that stops reading at the first
error. GenerateHandler had no cancel func at all, so one is added there.
Fixes#17825
Coding clients may insert runtime system messages after the initial user turn. The shared Qwen renderer rejected these transcripts before rendering, turning a potentially usable non-standard request into an HTTP 500.
Pass non-leading system turns through the existing raw ChatML path and warn when qwen3.8 encounters one. Extend the Anthropic tool-route integration scenario to cover this message pattern and remove the obsolete rejection test.
llama-server does not currently support WebP image payloads. Detect WebP media before forwarding, and transcode it to PNG. Pass all other media through unchanged.
Replace an existing vision integration image with a lossless WebP version so we now have coverage of JPG/PNG/WebP formats.
Fixes#17753
* qwen3.8: support developer instructions
Qwen3.8 does not define a developer role, while OpenAI-compatible coding agents commonly send developer instructions before user messages. Fold the leading system/developer instruction prefix into a single system turn before Qwen3.8 validation, preserving instruction precedence without changing Qwen3.5 or other renderer behavior.
Add streaming tool-call integration coverage for the native Ollama, OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages request shapes. Each case exercises prior assistant tool calls, tool results, follow-up rendering, and parsed tool-call output. Add Qwen3.8 to the release tools sweep.
Removes an unnecessary unit test that should not have been included in the original 3.8 PR.
* review comments