progress: join the render loop in stop and do the final writes after the
goroutine exits, so Stop/StopAndClear cannot race an in-flight render on
the shared bufio.Writer.
sched: read the unload-mutable fields in runnerRef.LogValue only under a
successful refMu.TryLock and omit them when contended, since slog resolves
it on goroutines that may already hold refMu.
The runner compiled a format as a JSON Schema, the one grammar kind its
xgrammar binding exposed. A structural tag holds a schema as one node of
a larger tree and also expresses what a schema cannot: free text around
constrained spans, a thinking region that closes before constrained
content, tool calls pinned to their schemas.
The runner now compiles structural tags only; its client wraps the API's
formats into one, which compiles to the same grammar as before. The JSON
token and vocabulary caps go with it: neither bounds compile cost, which
follows the grammar's state count. The byte and nesting caps stay.
A structured-output request could not use a model's draft head: it
decoded one token at a time, at roughly half the speculative throughput
on a dense 27B MTP model.
The grammar is enforced during verification instead: each draft
position's logits are masked before rejection sampling, so an invalid
draft is never accepted and every emitted token obeys the grammar.
Drafts stay unconstrained; constraining the draft chain would stall its
pipelined forwards.
Speculative steps also now dispatch the drafts before the host builds
the verification graph, worth 4-8% end to end at a fixed draft depth on
MTP models, with or without a grammar.
* ci: rebuild MLX macOS test payloads the release can't supply
The MLX unit test dependency payload can go stale when a PR changes the
xgrammar wrapper (release dylib lacks the new symbols) or bumps MLX pins (no
matching release, so MLX tests silently skipped). Compare the matched tag's
build rules and wrapper sources against the checkout, rebuilding only
libollama_xgrammar.dylib on drift; with no matching release, build one Metal
variant at the platform default.
* address comments
Safetensors gemma4 imports served by the MLX engine now answer image
and audio chats. Images run through both vision architectures: the
transformer tower (26B, 31B, e-series) and the 12B's encoder-free
unified embedder. Audio arrives through the same intake the ollama
API already accepts for gemma4 GGUFs — WAV bytes in the images field,
OpenAI input_audio parts, and /v1/audio/transcriptions uploads — with
the e2b/e4b checkpoints running clips through their conformer audio
encoder and the 12b unified checkpoint embedding the raw waveform
directly. Clips longer than 30 seconds are split evenly into chunks
of at most 30 seconds, cut at pauses, and encoded independently.
Each modality serves only checkpoints that carry it: 26B/31B have no
audio config and reject audio input, and checkpoints with an
unrecognized vision architecture still load as text-only models and
reject image requests.
The server previously hid the vision and audio capabilities for
gemma4 safetensors because the engine served neither. Both
suppressions are removed, and existing imports start advertising the
capabilities without re-importing since import already records them.
Audio-capable models need mono PCM at their expected sample rate
before model-specific feature extraction. Like the image decoder set
in base, the supported audio containers are decided once here so
every model accepts the same formats: WAV, covering integer and float
PCM, extensible headers, and multi-channel downmix. Anything else is
rejected as unrecognized; supporting another container later means
one new decoder here, with no model or runner changes.
Input at other sample rates is resampled through a band-limiting
filter, so mismatched rates degrade gracefully instead of aliasing.
Decoded clips are capped at ten minutes: the declared rate comes from
an untrusted header, and the cap is what keeps a small file claiming
an absurdly low rate from resampling into an enormous allocation.
Models whose encoder takes clips only up to a fixed length split
longer clips with Split. Chunks are sized evenly, and each cut moves
to the quietest point within a few seconds of its even share, so a
boundary lands on a pause where the clip has one instead of severing
a word between two independently encoded chunks. The even sizing
bounds the search so no chunk can end up over the limit.
The bindings captured MLX error messages but checked almost no calls,
so a failure continued with a null output and surfaced later as zero
results, skipped evals, or an unrelated crash.
Wrap every call in mlxCheck. Paths that return an error or disable a
GPU kernel backend use mlxError instead, and the lookups where a
non-zero status means a miss read the buffer first and then treat the
status as data.
Also free the string handles behind Array.String and the log values
after the call that fills them; they were freed before it and leaked.
MLX runs on one goroutine locked to its OS thread, so the thread-local
error buffers and closure-based check helpers defended against a
calling pattern that is already invalid.
Replace them with a single buffer that the handler fills and Go reads
after every call. mlxError returns the captured message; mlxCheck
panics on it and passes the call's result through, so a checked call
is one expression. Only an int status carries a failure signal, which
lets a message next to a zero status be reported as an earlier
unchecked call.
Fix two tests that relied on errors being dropped: the laguna
mixed-precision fixture used an unsupported quantization group size,
and the compile callback test expected the callback's own panic.
* Report cached prompt tokens
Add prompt_eval_cached_count to native responses and expose equivalent cached-token fields through the OpenAI- and Anthropic-compatible APIs. Keep prompt_eval_count as the logical input total while excluding cache hits from CLI and benchmark prefill rates. Surface processed and cached prompt counts in benchmark output.
Collect cache counts from llama-server and MLX, preserve coherent metrics across two-pass structured generation.
Fixes#8008
Related to #15758
* review comments
* ci: wire up MLX unit tests for PR runs
Download the latest Darwin release payload matching the current MLX and MLX-C revisions so macOS PR tests can exercise MLX without rebuilding it. If no matching release exists after a pin bump, leave MLX tests skipped until the next release.
Add whole-tree race coverage and smoke-run committed benchmarks. Verify generated UI types, and stabilize tests exposed by the broader CI coverage.
* review comments
* mlx: run tests on one pinned worker
Keep MLX tests and benchmarks on a shared pinned thread while preserving Fatal, Skip, and Cleanup semantics. Also clean stale CI payloads and ensure updater workers shut down cleanly.
* addres comments
* llama.cpp: version bump b10729
Regenerate the compat hooks patch for b10729: upstream removed the
whole-tensor load_data_for read (last consumer was llama-quantize,
which now reads slabs via load_data_range). Keep the existing hook
surface (constructor, skip loops, load_all_data, mtmd/clip) unchanged
and add maybe_load_text_tensor_range, which materializes a text load
op's output once per tensor and serves the new (offset, size) slab
reads from that cache.
* address comments
* Honor model generation defaults
Model-authored sampler defaults from GGUF metadata and HF generation_config.json were ignored, so built-in Ollama defaults could override model intent unless parameters were set in the Modelfile or request. The fix parses those defaults into model config and applies them before Modelfile/request options, preserving the expected precedence order.
* review comments
* address comments
* fix(docs): correct typos found during code review
Non-functional changes only:
- Fixed minor spelling mistakes in comments
- Corrected typos in user-facing strings
- No variables, logic, or functional code was modified.
Signed-off-by: Marcel Petrick <mail@marcelpetrick.it>
* fix additional typos and shell-unsafe example in docs
---------
Co-authored-by: Patrick Devine <patrick@ollama.com>
The MLX gemma3 port implements only the text stack, while gemma3 as GGUF
runs on llama-server with vision. Once MLX takes priority for
architectures both engines support, a registered gemma3 would route the
model to the engine that cannot serve images. No gemma3 safetensors
manifests were ever published, so removing the architecture affects no
existing installs and keeps gemma3 on llama-server.
Pi's Edit() only set baseUrl when creating a new ollama provider entry.
On subsequent launches it preserved whatever baseUrl was already in
~/.pi/agent/models.json, so switching OLLAMA_HOST to a remote server had
no effect — Pi would still connect to localhost.
Edit() now ensures baseUrl reflects the current OLLAMA_HOST. Models()
returns nil when the stored baseUrl no longer matches, so the launcher
only calls Edit() when the host has actually drifted. User-customized api
and apiKey fields are still preserved.
Model load code eagerly evaluated every weight fold (expert stacking,
gather transposes, gate/up fusing) as it was built, with the folds
running on the GPU against lazily loaded tensors: Metal committed
command buffers that waited on file reads, and macOS kills command
buffers that stall too long, so loading a large model from a slow
volume aborted with "Command buffer execution failed". The eager evals
also kept every layer's fold sources alive until the post-load sweep,
transiently holding roughly twice the expert weights on MoE models.
Build the folds lazily and let the runner's weight eval run them, and
on Metal materialize the loaded tensors with CPU reads before any
weight graph exists: no command buffer is ever committed waiting on
file data, at any storage speed, and fold sources free as their folds
execute. CUDA loads read at dispatch and skip the pre-pass. Models no
longer evaluate weights at load; on Metal, tensors the model does not
retain are now read before the sweep frees them.
Measured on an M5 Max, warm page cache, greedy outputs bit-identical:
before after
nemotron-3.5-lightning:30b-mlx 1.9s 39.7GiB 1.45s 24.7GiB
qwen3.6:35b-mlx 1.27s 22.5GiB 1.1-1.2s 22.4GiB
nemotron, reads at ~60MB/s aborts in 6s loads in 346s
Fixes#17902
The MLX runner accepted the API's format field but did not enforce it:
requests asking for JSON or a JSON Schema got unconstrained text, and
clients had no way to tell.
Enforce format with xgrammar: each sampling step masks the logits to
the tokens the grammar allows, so every emitted token and the end of
generation are valid under the constraint. Sampling, penalties, and
logprobs see the constrained distribution, and "json" yields a JSON
object, as the API documents and the llama-server path already
enforces. Only sampling waits on the mask; the forward pass is
dispatched before it, so constrained decoding stays pipelined.
The grammar engine is a dynamic library alongside MLX; when it is
missing, plain inference is unaffected and structured requests fail
with an explicit error. Constrained requests decode without
speculative decoding for now.
Decoding 256 tokens of a book-list schema on qwen3.8:27b-mlx (M5 Max,
seed 42, thinking off); pre-decode is the request time spent before
the first token:
unconstrained ~65 tok/s pre-decode ~70 ms
unconstrained, no draft ~32 tok/s pre-decode ~70 ms
JSON schema ~32 tok/s pre-decode ~70 ms
Schema and draft-less decoding are equal to within 0.1 tok/s in
paired adjacent requests, and a cold grammar compile adds nothing
measurable to pre-decode. The gap to unconstrained decoding is the
disabled draft model.
Fixes#16563
Co-authored-by: Daniel Hiltgen <daniel@ollama.com>
Token ids are int32 throughout the runner, so every caller reading ids
out of an int32 array narrowed the widened value right back. Make Int
and Ints return int32 and Float return float32, matching Floats, and
require the exact dtype instead of accepting and widening every
integer and float width: no caller read anything through those paths
but int32 tokens.
Ints and Floats also copied out of the array's buffer without
evaluating it first, so reading an array still in flight after an
async dispatch could return unwritten data, and correctness depended
on every call site remembering an explicit Eval. Evaluate in every
reader, matching the scalar readers, which already wait through item.
An available array costs a status check and an in-flight one waits
for its event; only a never-dispatched array evaluates a graph.
Nothing has set Grammar since the CGO engine removal took its writers
out; it survived as a read-only pass-through on the llama-server path
and a comment claiming it is set before dispatch. Remove the field and
the dead pass-through. llama-server keeps its wire-level grammar field,
which the "json" format conversion still uses.