Commit Graph
5688 Commits
Author SHA1 Message Date
Eva Ho 2a760bcbd0 app: make macOS update handoff retryable 2026-08-26 17:17:08 -04:00
Eva Ho 2cd22b8c0b app: fix macOS update process handoff 2026-08-26 10:51:53 -04:00
Parth Sareen e2c6c7e894 docs: claude docs (#18006) 2026-08-25 19:34:13 -07:00
Parth Sareen 9dbf139133 docs: document Claude Desktop integration (#18004) 2026-08-25 19:12:21 -07:00
Jesse Gross 77e3b0ac7a mlxrunner: avoid Metal GPU timeouts when loading models from slow storage
Model load code eagerly evaluated every weight fold (expert stacking,
gather transposes, gate/up fusing) as it was built, with the folds
running on the GPU against lazily loaded tensors: Metal committed
command buffers that waited on file reads, and macOS kills command
buffers that stall too long, so loading a large model from a slow
volume aborted with "Command buffer execution failed". The eager evals
also kept every layer's fold sources alive until the post-load sweep,
transiently holding roughly twice the expert weights on MoE models.

Build the folds lazily and let the runner's weight eval run them, and
on Metal materialize the loaded tensors with CPU reads before any
weight graph exists: no command buffer is ever committed waiting on
file data, at any storage speed, and fold sources free as their folds
execute. CUDA loads read at dispatch and skip the pre-pass. Models no
longer evaluate weights at load; on Metal, tensors the model does not
retain are now read before the sweep frees them.

Measured on an M5 Max, warm page cache, greedy outputs bit-identical:

                                    before           after
  nemotron-3.5-lightning:30b-mlx    1.9s  39.7GiB    1.45s    24.7GiB
  qwen3.6:35b-mlx                   1.27s 22.5GiB    1.1-1.2s 22.4GiB
  nemotron, reads at ~60MB/s        aborts in 6s     loads in 346s

Fixes #17902
2026-08-25 17:14:56 -07:00
Jesse Gross 147509c0c5 mlxrunner: add structured output support
The MLX runner accepted the API's format field but did not enforce it:
requests asking for JSON or a JSON Schema got unconstrained text, and
clients had no way to tell.

Enforce format with xgrammar: each sampling step masks the logits to
the tokens the grammar allows, so every emitted token and the end of
generation are valid under the constraint. Sampling, penalties, and
logprobs see the constrained distribution, and "json" yields a JSON
object, as the API documents and the llama-server path already
enforces. Only sampling waits on the mask; the forward pass is
dispatched before it, so constrained decoding stays pipelined.

The grammar engine is a dynamic library alongside MLX; when it is
missing, plain inference is unaffected and structured requests fail
with an explicit error. Constrained requests decode without
speculative decoding for now.

Decoding 256 tokens of a book-list schema on qwen3.8:27b-mlx (M5 Max,
seed 42, thinking off); pre-decode is the request time spent before
the first token:

    unconstrained              ~65 tok/s   pre-decode ~70 ms
    unconstrained, no draft    ~32 tok/s   pre-decode ~70 ms
    JSON schema                ~32 tok/s   pre-decode ~70 ms

Schema and draft-less decoding are equal to within 0.1 tok/s in
paired adjacent requests, and a cold grammar compile adds nothing
measurable to pre-decode. The gap to unconstrained decoding is the
disabled draft model.

Fixes #16563
Co-authored-by: Daniel Hiltgen <daniel@ollama.com>
2026-08-25 17:14:23 -07:00
Jesse Gross 7623501fc2 mlx: return exact types and evaluate arrays in the value readers
Token ids are int32 throughout the runner, so every caller reading ids
out of an int32 array narrowed the widened value right back. Make Int
and Ints return int32 and Float return float32, matching Floats, and
require the exact dtype instead of accepting and widening every
integer and float width: no caller read anything through those paths
but int32 tokens.

Ints and Floats also copied out of the array's buffer without
evaluating it first, so reading an array still in flight after an
async dispatch could return unwritten data, and correctness depended
on every call site remembering an explicit Eval. Evaluate in every
reader, matching the scalar readers, which already wait through item.
An available array costs a status check and an in-flight one waits
for its event; only a never-dispatched array evaluates a graph.
2026-08-25 17:14:23 -07:00
Jesse Gross 7027546ccf llm: remove the unused Grammar completion-request field
Nothing has set Grammar since the CGO engine removal took its writers
out; it survived as a read-only pass-through on the llama-server path
and a comment claiming it is set before dispatch. Remove the field and
the dead pass-through. llama-server keeps its wire-level grammar field,
which the "json" format conversion still uses.
2026-08-25 17:14:23 -07:00
Daniel Hiltgen 3d86a552b8 llama.cpp: version bump b10630 (#18003) 2026-08-25 17:11:26 -07:00
Daniel Hiltgen d465dc7ca1 MLX: version bump (#17955) 2026-08-25 17:11:13 -07:00
PhilippandCodex ad94d52965 cmake: make external compat patches idempotent (#17948)
Co-authored-by: Codex <noreply@openai.com>
2026-08-25 17:10:59 -07:00
Parth Sareen ebf200f952 proxy: preserve string content during image fallback (#18002) v0.33.0 v0.33.0-rc4 2026-08-25 15:03:57 -07:00
Eva H 075aa7e147 app: reset Claude Desktop models to defaults (#18000) 2026-08-25 17:43:21 -04:00
Eva H 377ef091dc app: keep Claude toggle busy while connecting (#17997) 2026-08-25 12:27:12 -07:00
Parth Sareen 6e19e916c7 app: reject browser origins on Claude Desktop gateway (#17989) 2026-08-25 09:48:44 -07:00
Eva H f6c59d8703 app: add Claude Desktop model mappings (#17979) v0.33.0-rc3 2026-08-24 22:43:47 -04:00
Parth Sareen 82ad9fa38b app: add Claude Desktop Auto mode setting (#17975) 2026-08-24 18:20:59 -07:00
Parth Sareen 60d83f8b0e app: make integrations list scrollable (#17977) 2026-08-24 17:46:45 -07:00
Eva H e2e82903fa app: improve desktop integration responsiveness (#17973)
* app: improve desktop integration responsiveness

* app: reconcile delayed Claude connection results

* app: preserve delayed Claude action errors
2026-08-24 19:25:24 -04:00
Eva H 939425152e app: fix desktop interaction regressions (#17970)
* app: fix desktop interaction regressions

* app: serialize settings reset updates
2026-08-24 19:13:48 -04:00
Anas Khan 02dc3ea4c3 cmd: guard empty editor before indexing fields (#17067)
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
2026-08-24 09:50:40 -07:00
Parth Sareen fb30760996 app: prevent Apps title bar overlap (#17925) v0.33.0-rc2 2026-08-21 18:48:22 -07:00
Devon Rifkin add1f92bdd launch: disable claude code token countdown to preserve KV cache (#17918)
Claude Code adds a "tokens left" system message after every tool
result. Since ollama moves system messages to the front of the prompt,
this breaks the KV cache on every request.
2026-08-21 17:44:44 -07:00
Parth Sareen 124e9af9d2 app: sign model recommendation endpoint (#17919) v0.33.0-rc1 2026-08-21 17:21:00 -07:00
Parth Sareen 2d9622a4d4 app: claude model management (#17915) v0.33.0-rc0 2026-08-21 15:04:58 -07:00
Eva H 30019c87c4 app: add Connect your apps experience (#17900) 2026-08-21 12:56:51 -07:00
Jesse Gross c44575ef14 mlxrunner: keep prefill snapshots when a request is cancelled mid-prompt
A long prompt records restore points during prefill, but they only
reached the prefix trie when the prefill completed; a cancelled request
closed and released everything it had captured. Agent clients routinely
cancel long prefills — their timeouts are shorter than the minutes a
40k-token prompt takes — so every retry started the whole prompt over
and never got further than the timeout allowed, which presents as the
model hanging forever.

Closing a session now attaches every snapshot the prefill crossed, so a
retry resumes from the last one and makes progress across timeouts.
Scenario tests cover retries resuming exactly where a cancelled attempt
stopped and cancellations on divergent conversation variants.

Fixes #17839
2026-08-21 09:58:33 -07:00
Jesse Gross 81f9a394e9 mlxrunner: grow the prefix trie by whole child nodes so restore points survive resumed prefills
A prefill that resumes partway into cached history — routine once
client timeouts interrupt long prompts — used to attach its captures
onto a node extended in place, so the stored snapshot spanned only the
tokens the prefill evaluated while the node's edge reached further
back. Restores walk node by node and trust each snapshot to cover its
node's edge; the short snapshot stranded the caches at mismatched
offsets and, on models with recurrent layers, ended up freeing all
cache state — a request matching 46k of a 47k-token prompt reprocessed
from zero.

Growth now never extends a node underneath its snapshots. New tokens
become a child node that carries exactly its own captures, and the
path stays compressed because non-user segments merge back into their
parent through the caches' snapshot Merge. Close already pages out
what it records, so every merge combines adjacent covered snapshots
and every stored snapshot spans exactly its node's edge.
2026-08-21 09:58:33 -07:00
Jesse Gross 30e2891808 mlxrunner: page out generated tokens when close records them
When a session closes, every cache rests exactly at the end of the
segment the trie is about to record. That is the one moment the
segment's state can be captured for every layer, so close now pages
the new segment out itself instead of recording it without snapshots
and leaving the capture to a later path switch.

Path switching then has nothing left to capture and only rewinds and
pages in. The whole-state entry taken at close is released when the
next request grows past the segment; sliding-window layers pay the
same window copy a scheduled capture already costs.
2026-08-21 09:58:33 -07:00
Jesse Gross b315b3ee97 mlxrunner: clip prefill captures to their trie node's edge
Page-in restores a path node by node and trusts each stored snapshot to
cover its node's whole edge. A capture taken during prefill spans from
the previous capture or the prefill base, which need not line up with
the node it lands on: when a prefill resumes partway into cached
history, a capture can reach back before its node's start, and a
capture landing on a node that already has snapshots replaced them
with a shorter span that page-in then could not serve.

Clip each capture to its node's edge on attach, and keep the snapshots
the node already has instead of replacing them.
2026-08-21 09:58:33 -07:00
Jesse Gross c01eafa552 mlxrunner: settle the draft caches when a prefill is cancelled
A prefill settles the drafter with the seed token after its last chunk,
leveling the draft caches with the targets; a cancelled prefill
returned before that, leaving the targets one token past the draft
caches and the recorded keys. The next request then had to move every
cache, and models with recurrent layers, which cannot rewind, fell back
to the last snapshot: a retry after a client timeout lost up to a full
snapshot interval of the prompt it had just evaluated.

Settle with the next prompt token on the cancelled path too. The caches
then rest level with the recorded keys, and a retry resumes exactly
where the prefill stopped.
2026-08-21 09:58:33 -07:00
Parth Sareen 8f912415e8 launch: fall back to npx for DeepSeek Harness (#17758) 2026-08-20 14:04:44 -07:00
Eva H 5ad1681cf1 polish onboarding layout and disable zoom (#17885) 2026-08-20 17:03:58 -04:00
Parth Sareen 30546d1fd4 app: add claude desktop app (#17899) 2026-08-20 14:03:43 -07:00
Daniel Hiltgen 6bba484f1a lint fixes (#17897) 2026-08-20 10:23:25 -07:00
Daniel Hiltgen e92b7855f6 mlx update (#17886) 2026-08-20 10:02:41 -07:00
Daniel Hiltgen 4e13421378 mlx: fix mac assumptions on linux/windows (#17898)
The default packaging was broken due to mac assumptions
leaking into windows
2026-08-20 09:50:10 -07:00
Eva H b7871fc0d1 app: add desktop onboarding flow (#17853) v0.32.15-rc2 v0.32.15 2026-08-19 15:40:16 -07:00
Daniel Hiltgen e0c95a5ffd server: don't wedge chat and generate on a mid-stream parser error (#17883)
When a builtin parser rejects model output, the completion callback wrote the
error to an unbuffered channel and returned. The callback cannot stop
generation -- it has no error return -- so the next chunk re-entered the
callback, hit the same parse error and blocked writing to a channel the
consumer had already stopped reading after emitting its 500. The completion
never returned, the goroutine leaked and the runner request was never
released, so retrying the same prompt hung with no log output until the client
gave up.

Record the parse error, cancel the completion, and report it once the
completion has returned. Parse failures landing on the final chunk were
already terminal, which is why non-thinking requests and the direct
qwen3-coder parser path failed cleanly and only thinking mode wedged.

ChatHandler and GenerateHandler share the defect: both run the same parser in
the same shape of callback behind a consumer that stops reading at the first
error. GenerateHandler had no cancel func at all, so one is added there.

Fixes #17825
2026-08-19 15:14:50 -07:00
Parth Sareen b8a6272440 qwen3.8: normalize system messages (#17855) 2026-08-19 13:08:09 -07:00
Daniel Hiltgen d1bd15ccce ci: plumb temporary MLX patch through to docker stages (#17874)
Follow up to #17850
v0.32.15-rc1
2026-08-19 08:33:53 -07:00
Daniel Hiltgen 0bb0925920 mlx update (#17850)
Temporarily carry https://github.com/ml-explore/mlx-c/pull/127
v0.32.15-rc0
2026-08-19 07:15:11 -07:00
Gaurav Garg a5165c53ac Add a model metadata cache to reduce Ollama’s per-request overhead (#17752) 2026-08-18 12:22:51 -07:00
Daniel Hiltgen cd37044093 llama.cpp update (#17851) 2026-08-18 11:53:00 -07:00
Daniel Hiltgen d67ad83426 mlx update (#17761) v0.32.14 v0.32.14-rc0 2026-08-15 11:56:40 -07:00
Daniel Hiltgen e5a81899d0 llama.cpp update (#17760) 2026-08-14 19:03:20 -07:00
Parth Sareen 78e818e3ce docs: register DeepSeek Harness (#17751) 2026-08-14 14:30:21 -07:00
Daniel Hiltgen 87abaa019e renderers/qwen: tolerate non-leading system messages (#17757)
Coding clients may insert runtime system messages after the initial user turn. The shared Qwen renderer rejected these transcripts before rendering, turning a potentially usable non-standard request into an HTTP 500.

Pass non-leading system turns through the existing raw ChatML path and warn when qwen3.8 encounters one. Extend the Anthropic tool-route integration scenario to cover this message pattern and remove the obsolete rejection test.
2026-08-14 14:12:50 -07:00
Daniel Hiltgen f427fa0753 llm: transcode WebP images for llama-server (#17755)
llama-server does not currently support WebP image payloads. Detect WebP media before forwarding, and transcode it to PNG. Pass all other media through unchanged.

Replace an existing vision integration image with a lossless WebP version so we now have coverage of JPG/PNG/WebP formats.

Fixes #17753
2026-08-14 13:21:11 -07:00
Daniel Hiltgen 0f25c31bd5 qwen3.8: support developer instructions (#17749)
* qwen3.8: support developer instructions

Qwen3.8 does not define a developer role, while OpenAI-compatible coding agents commonly send developer instructions before user messages. Fold the leading system/developer instruction prefix into a single system turn before Qwen3.8 validation, preserving instruction precedence without changing Qwen3.5 or other renderer behavior.

Add streaming tool-call integration coverage for the native Ollama, OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages request shapes. Each case exercises prior assistant tool calls, tool results, follow-up rendering, and parsed tool-call output. Add Qwen3.8 to the release tools sweep.

Removes an unnecessary unit test that should not have been included in the original 3.8 PR.

* review comments
v0.32.13
2026-08-14 11:30:27 -07:00