5777 Commits
Author SHA1 Message Date
Parth Sareen ebf200f952 proxy: preserve string content during image fallback (#18002) v0.33.0 v0.33.0-rc4 2026-08-25 15:03:57 -07:00
Eva H 075aa7e147 app: reset Claude Desktop models to defaults (#18000) 2026-08-25 17:43:21 -04:00
Eva H 377ef091dc app: keep Claude toggle busy while connecting (#17997) 2026-08-25 12:27:12 -07:00
Parth Sareen 6e19e916c7 app: reject browser origins on Claude Desktop gateway (#17989) 2026-08-25 09:48:44 -07:00
Eva H f6c59d8703 app: add Claude Desktop model mappings (#17979) v0.33.0-rc3 2026-08-24 22:43:47 -04:00
Parth Sareen 82ad9fa38b app: add Claude Desktop Auto mode setting (#17975) 2026-08-24 18:20:59 -07:00
Parth Sareen 60d83f8b0e app: make integrations list scrollable (#17977) 2026-08-24 17:46:45 -07:00
Eva H e2e82903fa app: improve desktop integration responsiveness (#17973)
* app: improve desktop integration responsiveness

* app: reconcile delayed Claude connection results

* app: preserve delayed Claude action errors
2026-08-24 19:25:24 -04:00
Eva H 939425152e app: fix desktop interaction regressions (#17970)
* app: fix desktop interaction regressions

* app: serialize settings reset updates
2026-08-24 19:13:48 -04:00
Anas Khan 02dc3ea4c3 cmd: guard empty editor before indexing fields (#17067)
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
2026-08-24 09:50:40 -07:00
Parth Sareen fb30760996 app: prevent Apps title bar overlap (#17925) v0.33.0-rc2 2026-08-21 18:48:22 -07:00
Devon Rifkin add1f92bdd launch: disable claude code token countdown to preserve KV cache (#17918)
Claude Code adds a "tokens left" system message after every tool
result. Since ollama moves system messages to the front of the prompt,
this breaks the KV cache on every request.
2026-08-21 17:44:44 -07:00
Parth Sareen 124e9af9d2 app: sign model recommendation endpoint (#17919) v0.33.0-rc1 2026-08-21 17:21:00 -07:00
Parth Sareen 2d9622a4d4 app: claude model management (#17915) v0.33.0-rc0 2026-08-21 15:04:58 -07:00
Eva H 30019c87c4 app: add Connect your apps experience (#17900) 2026-08-21 12:56:51 -07:00
Jesse Gross c44575ef14 mlxrunner: keep prefill snapshots when a request is cancelled mid-prompt
A long prompt records restore points during prefill, but they only
reached the prefix trie when the prefill completed; a cancelled request
closed and released everything it had captured. Agent clients routinely
cancel long prefills — their timeouts are shorter than the minutes a
40k-token prompt takes — so every retry started the whole prompt over
and never got further than the timeout allowed, which presents as the
model hanging forever.

Closing a session now attaches every snapshot the prefill crossed, so a
retry resumes from the last one and makes progress across timeouts.
Scenario tests cover retries resuming exactly where a cancelled attempt
stopped and cancellations on divergent conversation variants.

Fixes #17839
2026-08-21 09:58:33 -07:00
Jesse Gross 81f9a394e9 mlxrunner: grow the prefix trie by whole child nodes so restore points survive resumed prefills
A prefill that resumes partway into cached history — routine once
client timeouts interrupt long prompts — used to attach its captures
onto a node extended in place, so the stored snapshot spanned only the
tokens the prefill evaluated while the node's edge reached further
back. Restores walk node by node and trust each snapshot to cover its
node's edge; the short snapshot stranded the caches at mismatched
offsets and, on models with recurrent layers, ended up freeing all
cache state — a request matching 46k of a 47k-token prompt reprocessed
from zero.

Growth now never extends a node underneath its snapshots. New tokens
become a child node that carries exactly its own captures, and the
path stays compressed because non-user segments merge back into their
parent through the caches' snapshot Merge. Close already pages out
what it records, so every merge combines adjacent covered snapshots
and every stored snapshot spans exactly its node's edge.
2026-08-21 09:58:33 -07:00
Jesse Gross 30e2891808 mlxrunner: page out generated tokens when close records them
When a session closes, every cache rests exactly at the end of the
segment the trie is about to record. That is the one moment the
segment's state can be captured for every layer, so close now pages
the new segment out itself instead of recording it without snapshots
and leaving the capture to a later path switch.

Path switching then has nothing left to capture and only rewinds and
pages in. The whole-state entry taken at close is released when the
next request grows past the segment; sliding-window layers pay the
same window copy a scheduled capture already costs.
2026-08-21 09:58:33 -07:00
Jesse Gross b315b3ee97 mlxrunner: clip prefill captures to their trie node's edge
Page-in restores a path node by node and trusts each stored snapshot to
cover its node's whole edge. A capture taken during prefill spans from
the previous capture or the prefill base, which need not line up with
the node it lands on: when a prefill resumes partway into cached
history, a capture can reach back before its node's start, and a
capture landing on a node that already has snapshots replaced them
with a shorter span that page-in then could not serve.

Clip each capture to its node's edge on attach, and keep the snapshots
the node already has instead of replacing them.
2026-08-21 09:58:33 -07:00
Jesse Gross c01eafa552 mlxrunner: settle the draft caches when a prefill is cancelled
A prefill settles the drafter with the seed token after its last chunk,
leveling the draft caches with the targets; a cancelled prefill
returned before that, leaving the targets one token past the draft
caches and the recorded keys. The next request then had to move every
cache, and models with recurrent layers, which cannot rewind, fell back
to the last snapshot: a retry after a client timeout lost up to a full
snapshot interval of the prompt it had just evaluated.

Settle with the next prompt token on the cancelled path too. The caches
then rest level with the recorded keys, and a retry resumes exactly
where the prefill stopped.
2026-08-21 09:58:33 -07:00
Parth Sareen 8f912415e8 launch: fall back to npx for DeepSeek Harness (#17758) 2026-08-20 14:04:44 -07:00
Eva H 5ad1681cf1 polish onboarding layout and disable zoom (#17885) 2026-08-20 17:03:58 -04:00
Parth Sareen 30546d1fd4 app: add claude desktop app (#17899) 2026-08-20 14:03:43 -07:00
Daniel Hiltgen 6bba484f1a lint fixes (#17897) 2026-08-20 10:23:25 -07:00
Daniel Hiltgen e92b7855f6 mlx update (#17886) 2026-08-20 10:02:41 -07:00
Daniel Hiltgen 4e13421378 mlx: fix mac assumptions on linux/windows (#17898)
The default packaging was broken due to mac assumptions
leaking into windows
2026-08-20 09:50:10 -07:00
Eva H b7871fc0d1 app: add desktop onboarding flow (#17853) v0.32.15-rc2 v0.32.15 2026-08-19 15:40:16 -07:00
Daniel Hiltgen e0c95a5ffd server: don't wedge chat and generate on a mid-stream parser error (#17883)
When a builtin parser rejects model output, the completion callback wrote the
error to an unbuffered channel and returned. The callback cannot stop
generation -- it has no error return -- so the next chunk re-entered the
callback, hit the same parse error and blocked writing to a channel the
consumer had already stopped reading after emitting its 500. The completion
never returned, the goroutine leaked and the runner request was never
released, so retrying the same prompt hung with no log output until the client
gave up.

Record the parse error, cancel the completion, and report it once the
completion has returned. Parse failures landing on the final chunk were
already terminal, which is why non-thinking requests and the direct
qwen3-coder parser path failed cleanly and only thinking mode wedged.

ChatHandler and GenerateHandler share the defect: both run the same parser in
the same shape of callback behind a consumer that stops reading at the first
error. GenerateHandler had no cancel func at all, so one is added there.

Fixes #17825
2026-08-19 15:14:50 -07:00
Parth Sareen b8a6272440 qwen3.8: normalize system messages (#17855) 2026-08-19 13:08:09 -07:00
Daniel Hiltgen d1bd15ccce ci: plumb temporary MLX patch through to docker stages (#17874)
Follow up to #17850
v0.32.15-rc1
2026-08-19 08:33:53 -07:00
Daniel Hiltgen 0bb0925920 mlx update (#17850)
Temporarily carry https://github.com/ml-explore/mlx-c/pull/127
v0.32.15-rc0
2026-08-19 07:15:11 -07:00
Gaurav Garg a5165c53ac Add a model metadata cache to reduce Ollama’s per-request overhead (#17752) 2026-08-18 12:22:51 -07:00
Daniel Hiltgen cd37044093 llama.cpp update (#17851) 2026-08-18 11:53:00 -07:00
Daniel Hiltgen d67ad83426 mlx update (#17761) v0.32.14 v0.32.14-rc0 2026-08-15 11:56:40 -07:00
Daniel Hiltgen e5a81899d0 llama.cpp update (#17760) 2026-08-14 19:03:20 -07:00
Parth Sareen 78e818e3ce docs: register DeepSeek Harness (#17751) 2026-08-14 14:30:21 -07:00
Daniel Hiltgen 87abaa019e renderers/qwen: tolerate non-leading system messages (#17757)
Coding clients may insert runtime system messages after the initial user turn. The shared Qwen renderer rejected these transcripts before rendering, turning a potentially usable non-standard request into an HTTP 500.

Pass non-leading system turns through the existing raw ChatML path and warn when qwen3.8 encounters one. Extend the Anthropic tool-route integration scenario to cover this message pattern and remove the obsolete rejection test.
2026-08-14 14:12:50 -07:00
Daniel Hiltgen f427fa0753 llm: transcode WebP images for llama-server (#17755)
llama-server does not currently support WebP image payloads. Detect WebP media before forwarding, and transcode it to PNG. Pass all other media through unchanged.

Replace an existing vision integration image with a lossless WebP version so we now have coverage of JPG/PNG/WebP formats.

Fixes #17753
2026-08-14 13:21:11 -07:00
Daniel Hiltgen 0f25c31bd5 qwen3.8: support developer instructions (#17749)
* qwen3.8: support developer instructions

Qwen3.8 does not define a developer role, while OpenAI-compatible coding agents commonly send developer instructions before user messages. Fold the leading system/developer instruction prefix into a single system turn before Qwen3.8 validation, preserving instruction precedence without changing Qwen3.5 or other renderer behavior.

Add streaming tool-call integration coverage for the native Ollama, OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages request shapes. Each case exercises prior assistant tool calls, tool results, follow-up rendering, and parsed tool-call output. Add Qwen3.8 to the release tools sweep.

Removes an unnecessary unit test that should not have been included in the original 3.8 PR.

* review comments
v0.32.13
2026-08-14 11:30:27 -07:00
Daniel Hiltgen 5512797527 qwen3.8: add renderer and MLX import support (#17745)
Qwen3.8 keeps the Qwen3.5 model architecture and parser, but its chat template adds reasoning-effort and preserved-thinking semantics. Detect those template markers during safetensors import, select the qwen3.8 renderer, and cover thinking, tools, continuation, and malformed parser input.

Make indexed safetensors imports use the weight map's shard names instead of independently filtering files by the model-* convention. Reject unsafe shard paths, ignore unindexed tensors, and fail when an indexed weight is missing or stored in a different shard. Retain the conservative model-* scan when no index is present.

Treat Classification.Quantize as the effective tensor format and pass it to the manifest writer. This records file_type for automatic block-FP8-to-MXFP8 conversion and recognized prequantized inputs, preserves requested quantization and base-plus-draft behavior, and avoids claiming one type for mixed or unknown formats.

Normalize both supported convolution weight layouts with an explicit reshape. Add focused unit coverage for renderer selection, parser behavior, shard inventory, manifest metadata, and convolution layout; heavyweight reference-forward and release integration checks remain bring-up artifacts.
2026-08-14 09:31:09 -07:00
Parth Sareen 39df91c982 launch: add DeepSeek Harness integration (#17733) v0.32.11 2026-08-13 15:19:02 -07:00
Daniel Hiltgen 7ce88bd686 model/renderers: match Muse Glimmer reasoning template (#17732)
Updates Muse Glimmer Jinja reference template to the latest publisher version and mirror its explicit-system reasoning handling in the Go renderer.

Explicit system prompts now normalize "Reasoning effort" to "Reasoning strength" and skip adding a renderer-provided reasoning line when the prompt already contains one. This prevents duplicate or conflicting reasoning directives while preserving the default-system behavior.

Add reference tests for both normalization and deduplication, including Jinja-backed validation.
2026-08-13 14:35:50 -07:00
Daniel Hiltgen 01d04d50f8 launch: add Muse Code integration (#17594)
* launch: add Muse Code integration

Add `ollama launch muse` for Meta's Muse Code CLI.

Muse only takes a model catalog from settings.json (normally it fetches one from its provider and refuses to start otherwise), and that file's endpoint_transport is a global provider switch. So the integration writes a settings file under its own config root (~/.ollama/launch/muse-config via XDG_CONFIG_HOME), leaving a Meta-backed muse install untouched, and re-seeds it from muse's own persisted copy on later runs.

The launched model is preloaded so its catalog row carries the context length the server actually allocated, not the trained maximum; the loaded-context helpers move from cmd/agent_tui.go into cmd/launch for reuse.

Muse sends reasoning efforts outside Ollama's scale (minimal, xhigh, ultra), which were hard 400s; clamp them to the nearest tier in one helper shared by the chat and responses converters.

The registry entry stays Hidden (alias "muse-code"), like kimi and vscode.

* review comments

* skip muse test on windows (unsupported platform)
2026-08-13 13:10:29 -07:00
Parth Sareen 9a56a0e845 agent: allow multiple edits per edit tool call (#17711) v0.32.10 2026-08-12 16:45:53 -07:00
Daniel Hiltgen 88313499e0 mlx: avoid pulling MLX models when MLX is missing (#17710)
As we look to bring Linux and Windows MLX support online, instead of blocking
downloads at the registry to avoid users wasting time downloading a model they
can't run, shift the logic to the local side which knows if MLX is present or not.
v0.32.10-rc1
2026-08-12 14:42:17 -07:00
Jesse Gross 2b4a99376c nn: speed up prefill on double-scale nvfp4 models
ModelOpt checkpoints apply a float32 global scale to every projection
output on top of the per-group quantization scales. Running the
multiply and the cast back to the activation dtype as separate eager
ops costs an extra kernel launch and a materialized intermediate per
projection.

Compile the multiply and cast into one kernel. On an M5 Max (medians
of order-swapped A/B runs against main; greedy outputs byte-identical):

    qwen3.6:27b        prefill  703 -> 769 t/s  +7.9%
    muse-glimmer:30b   prefill  790 -> 843 t/s  +6.7%

Speculative decode is unchanged within noise on both models. Only
checkpoints with a global scale are affected; single-scale nvfp4,
mxfp8, and affine checkpoints take the unchanged path.
v0.32.10-rc0
2026-08-12 13:25:33 -07:00
Daniel Hiltgen e922bc7125 llama.cpp bump (#17702) 2026-08-12 12:10:18 -07:00
Daniel Hiltgen 950dd9ac67 MLX update (#17704) 2026-08-12 12:09:50 -07:00
Parth Sareen b6b1b258c3 openai: support web search in Responses API (#17686) 2026-08-12 11:51:54 -07:00
VigneshandPatrick Devine 4138e853d5 server/images: prevent skipVerify map collision with duplicate digests (#15504)
When a manifest contains a config and layer with the same digest, the
skipVerify map entry was overwritten by the config's cache-hit value
(true), replacing the layer's non-cache-hit value (false). This caused
verifyBlob to be skipped for the freshly downloaded blob.

A rogue OCI registry could exploit this by serving a manifest with
duplicate digests and redirecting blob downloads to internal endpoints.
The SSRF response would be written to disk, hash verification would be
skipped due to the map collision, and the blob would persist.

The fix uses logical AND when updating skipVerify: once any download of
a digest was not a cache hit, verification is always performed.

Fixes #15485

---------

Co-authored-by: Patrick Devine <patrick@ollama.com>
2026-08-12 11:44:32 -07:00