Listing, pulling, creating, copying and deleting models from the Manage Ollama dialog ignored the connection's custom headers and authentication type, so the dialog failed behind gateways such as Cloudflare Access and sent the key as a Bearer token even with the authentication type set to None. Checking a single connection's version sent no key at all. All of these, and the other requests to an Ollama connection such as text generation and embeddings, now use the connection's headers and authentication type, matching what verifying the connection and chatting already do.
Fixes#31487
Choosing Share or Delete from a chat's menu in the sidebar left the menu open behind the dialog, so the first click inside the dialog only closed the menu behind it. Copy Link and Confirm had to be clicked twice. Share and Delete now close the menu before opening their dialog, like Rename, Pin and Clone already do.
Fixes#31486
With native function calling, citations produced by query_knowledge_files and query_chat_files never showed the relevance percentage badge, while the same knowledge base queried through classic RAG did.
The tools already return a distance per chunk, but the step that groups tool results into citation sources dropped it. Each grouped source now carries a distances list aligned with its documents, the same shape the classic RAG path emits, so the existing citation UI shows the badge without frontend changes. Chunks without a score (notes) leave the list empty, which the UI already treats as no score.
Fixes#29776
When the provider failed partway through a streaming request to the Anthropic Messages endpoint, the stream still ended like a normally finished answer, so Claude Code and the Anthropic SDKs took the cut-off text as complete. The stream now ends with an Anthropic error event, with the provider's error message if it sent one, so clients raise an error. Successful streams are unchanged.
Fixes#31403
Tavily web search always ran at Tavily's default depth (basic), because the search request never sent `search_depth`. The only Tavily depth control in Admin > Settings > Web Search, "Tavily Extract Depth", applies to the Extract API used by the web loader, never to search.
This adds `TAVILY_SEARCH_DEPTH` (env var and persisted setting, default `basic`) and a "Tavily Search Depth" select (ultra-fast, fast, basic, advanced) under the Tavily search engine settings. The value is sent as `search_depth` on every Tavily search request, so admins can set search and extract depth independently, for example fast search with advanced extraction.
The default matches Tavily's own default, so existing setups keep the same behaviour until the setting is changed.
Fixes#29891
Saving an arena model in Admin Settings > Models (for example to set default tools or capabilities) creates a model entry with the arena id. That entry replaced the arena model's metadata wholesale, dropping the access grants, model_ids and filter_mode configured in Admin Settings > Evaluations. From then on every non-admin user lost the arena model, even when it was public, and chats through it ignored the configured model pool.
The override now keeps those three keys from the evaluation config, which is where arena access and the model pool are managed. Everything else set in Settings > Models (tools, capabilities, description, profile image) still applies.
Verified end to end on base and patched: after the override a user sees and can chat with a public arena model (base: hidden, 400), private arena models stay hidden, and 20 admin chats all route to the configured pool (base: spread across all models).
Fixes#29564
In "Ask for approval" mode, when the model requested several tools in one turn, only the first call got an approval card. The others stayed on "Executing..." forever, never ran, could not be approved (the server answered "already resolved"), and the model was called again without their results. The stuck state was saved to the chat.
Once streaming finishes, every call in the turn is marked as completed (arguments done, nothing run yet). The approval pause only queued siblings that were still in progress, so these were skipped. They are now queued as well, and each one gets its own approval card in turn after the previous one is resolved.
Calls that already have a result and rejected calls are untouched, and single-call turns behave as before. Verified against the real approval functions with same-name, mixed-name, reject and ask_user batches, plus the tests-repo unit suite (identical results before and after).
Fixes#29293
Anthropic clients such as Claude Code send "tools": [] on text-only requests like prompt-hook evaluation. The Anthropic Messages endpoint carried that empty array into the converted OpenAI request, and vLLM and the OpenAI API reject it with HTTP 400, so those requests failed while normal chats with tools kept working. A "tools": null body crashed the converter with a 500.
The converter now only emits tools when the list is non-empty, and only emits tool_choice when tools were emitted. Dropping tools alone is not enough: the same backends also reject tool_choice without tools, so a request sending an empty tool list plus a tool_choice would still fail.
Requests with real tools are converted exactly as before. Verified end to end against a mock backend enforcing vLLM's validation: empty, null and tool_choice-only requests went from 400/500 to 200 with end_turn, streaming included.
Fixes#31341
Bumps pillow 12.2.0 -> 12.3.0 and aiohttp 3.13.5 -> 3.14.3 in requirements.txt, requirements-slim.txt and pyproject.toml. Both are upstream security releases.
uv.lock is regenerated, which also brings it in line with the hiredis 3.4.2 pin from #31328.
Verified on a real install of the bumped set: the backend boots with /health 200, the pillow and aiohttp contract tests pass (84/84), and the full tests suite shows the same results as on dev (no new failures).
Attaching a link in chat or to a knowledge base that cannot be fetched (closed port, blocked by the fetch filter, an HTTP error such as 404) showed the toast "Error processing URL", which never said which link failed or that fetching it was the problem.
The fetch step now answers with "Could not read content from <url>", the same message process/web gives for a link it cannot read, so both endpoints report a dead link the same way. The too-large 413 still passes through unchanged, and a working link returns exactly what it did before.
The new handler covers only the fetch. Rewording the endpoint's existing catch-all would be one line, but that handler also receives database errors from the config and file lookups, which would then be reported as an unreadable link.
Related to #31347
With a personal tool server connection such as Open Terminal, the main model could call its tools but a foreground sub-agent it delegated to got none of them. Chats resuming after a tool approval lost those tools the same way. Setting up the tools for the main model emptied the list those later steps read from. It now works on a copy, so sub-agents and resumed chats get the same tools as the parent.
Fixes#29893
Before a share link exists, the Share dialog tells users that anyone with the URL will be able to view the chat. A new link is actually private: only its owner and admins can open it until the access is changed to Public or Open, or specific users or groups are added. The text now says the link stays private until you choose who can view it, which matches how sharing works and what the docs describe.
Fixes#31417
On PostgreSQL with ENABLE_ORJSON off (the default), searching automations by a word from their prompt, or filtering models by a tag, found nothing when the word was non-ASCII, for example Chinese. SQLite, and PostgreSQL with ENABLE_ORJSON on, were fine. With the default setting non-ASCII text is saved as \uXXXX codes, and PostgreSQL reads the backslash in a search pattern as a special character, so the search never matched. Special characters in the search text are now taken literally, so these searches work on both databases, and a % or _ typed into them now matches only itself.
Fixes#31422
With the default content extraction engine, backslashes in an uploaded .html or .htm file were read as escape sequences. A path like C:\new\table was saved with a line break and a tab in it, and a page containing C:\Users failed to upload with a 'unicodeescape' codec error. HTML files are now read the same way as .txt and .md uploads, so the saved text matches the page.
Fixes#31440
With hybrid search on and a vector database without built-in hybrid search, a collection that failed to load (for example Qdrant strict mode rejecting the request) was skipped quietly, so retrieval returned no documents and never fell back to normal vector search. A failed load now counts as a failed collection, so when every collection fails retrieval falls back to vector search, the same way it already does when the search itself fails. The retrieval API returns its usual error in that case.
Part of #31459
With Qdrant strict mode on and a max_query_limit below 999999999, Qdrant rejects Open WebUI's reads with "Limit exceeded", so every file upload after the first fails in the default multitenancy mode, and hybrid search finds nothing. Reads now go in pages of 1000 points, so any strict-mode limit of 1000 or more works. Without strict mode the results are the same as before.
Tested against Qdrant 1.19.1 with max_query_limit 1000, for both multitenancy on and off: collections of up to 2500 points come back complete, limits are respected, and tenants stay separated.
Fixes#31459
Fills all 191 strings that were still untranslated in the German catalog and fixes existing entries that turned up while checking each string against the spot where it appears in the interface. Several were plainly wrong: "Reason" in the feedback dialog read "Nachdenken" (thinking), the Open Terminal product name was translated as the action "open terminal", the terminal memory limit said "Erinnerungen" (memories) and the share dialog assembled into broken sentences. Leftover English ("API Key", "History", "Policy", "Queue"), informal "du" toward the user and outliers like "Nutzer" next to the dominant "Benutzer" now match the rest of the catalog, and compounds written as two words are joined or hyphenated. Prompts written to the AI model (suggestion cards, system prompt examples) use the informal "du", and the chat bubble label stays "Du". Only de-DE is touched, with no keys added, removed or reordered.
When a scheduled timer failed before the model started answering, for example because its model had been removed, the chat showed the timer's prompt with a blank reply that looked stuck, and the error never appeared in the chat. The reply now shows the error and stops loading, like any other failed message.
Fixes#31481
A link whose host refuses the connection, such as a closed port, still came back from POST /api/v1/retrieval/process/web as "Error querying knowledge base", so the caller was told the knowledge base failed when the link was the problem.
The web loaders log a failed fetch and return no documents. That empty result then failed while being saved to the vector store, and the save error was the one reported.
process_web now answers with the existing "Could not read content from <url>" 400 as soon as the loader returns no documents, the same message a link refused by the fetch filter already gets. The check sits in the endpoint so web search and the other users of the loaders keep their current behaviour.
With process=false or embedding bypassed, an unreachable link now gets the same 400 where it used to return 200 with empty content.
Fixes#31347
Picking a skill from the / menu in the chat input inserted it as an @ mention, so the skill was never loaded and its raw tag ended up in the message the model received. Picking the same skill from the $ menu worked. A skill chosen through / is now sent the same way as through $ and gets applied.
Fixes#29978
When a model wraps its answer in <|begin_of_solution|> and <|end_of_solution|>, only the opening marker was removed. The closing marker stayed visible in the reply and was saved with the message, and anything the model wrote after it was glued onto the answer. Now both markers are removed and text after the answer shows up as a normal part of the reply.
Fixes#31434
Sometimes a reply where the model used tools gets saved with a tool result that no call in that reply asked for, or with a tool call that never got its result. The chat history was then sent to the provider unchanged, Anthropic and Bedrock rejected it, and every following message in that chat failed until the user deleted the broken reply. Now each tool call is only kept together with its own result from the same reply, and the unmatched calls and results are left out of what gets sent to the model. The chat itself is not changed, and correctly saved chats are sent exactly as before.
Fixes#28937
When an API request to an Ollama model set max_tokens, Open WebUI passed it on in a place Ollama does not read, so Ollama ignored it and replies ran to full length. The limit now reaches Ollama as its own output length setting, so replies stop at the requested length. It also wins over a max_tokens value saved in the model's advanced parameters, as the API docs describe. Chats in the web UI were not affected, since their limit already reached Ollama correctly.
Fixes#31432
With reasoning models, when the first part of the answer arrived together with the end of the thinking block and ended with a space, that space went missing, so "The answer is 4." was shown and saved as "The answeris 4.". That space is now kept.
Fixes#31435
With a connection set to the Responses API, when the provider reported a reply as failed, the error showed while streaming but was gone after a reload, leaving an empty reply. Some other provider errors never showed up at all, not even while streaming. Both kinds of error now show up and are still there after a reload, the same as on Chat Completions connections.
Fixes#31433
On the default SQLite setup (session sharing off), a POST to /api/v1/models/sync containing any model that already exists answered 200 with an empty list and stored nothing. After a 5 second stall, the only trace was "database is locked" in the server log.
The sync wrote each model's access grants while the model update was still uncommitted. Without session sharing the grant writes run on a second database session, and SQLite allows one writer at a time, so that write waited on the same request's uncommitted update until the busy timeout expired and the whole sync was dropped.
Grants are now written after the model changes are committed, the same order model create and update already use. A failed model commit now also leaves every grant untouched. PostgreSQL and setups with session sharing on behave as before.
Fixes#31346
With ENABLE_OAUTH_GROUP_CREATION on, every group in a user's OAuth claim was created on login, including groups matching OAUTH_BLOCKED_GROUPS. Membership sync already ignored those groups, so the result was empty groups nobody could join. With IdPs that send a user's full directory membership (Keycloak backed by LDAP/AD), one login could fill the group table with thousands of them.
Group creation now applies the same blocklist check as the membership add and remove steps, so a blocked group is never created, joined or left through OAuth. Groups that are not blocked are created as before.
Fixes#29558
Korean EUC-KR and Japanese Shift-JIS text files were stored as garbled Chinese characters, so retrieval, knowledge bases and the model context all worked on text that is not in the file.
Encoding detection puts chardet's guess in front of a fixed GB18030, Big5, EUC-KR, EUC-JP try order, and GB18030 decodes almost any double-byte text without an error. The guess map was written for chardet 5. Since the bump to chardet 7 in v0.10.0, Korean text is reported as CP949 and Japanese text as cp932 or SHIFT_JIS, which the map either did not know or dropped because the codec was not in the try order, so these files fell through to GB18030.
The map now covers CP949 and cp932, and a mapped guess is always tried first. SHIFT_JIS maps to cp932, the Windows superset, because chardet also reports SHIFT_JIS for ordinary Japanese files containing characters such as ① or ㈱ that plain Shift-JIS cannot decode; this is the same subset-to-superset rule the map already applies to GB2312.
Korean and Japanese files now decode correctly, and Chinese, EUC-JP, UTF-8 and Western files decode as before. The one trade-off of trusting the guess: chardet 7 labels some files holding only a few Chinese characters (a short label or a one-line comment) as CP949, and those now read as Korean. No regressions were found in files with more Chinese text than that.
Fixes#31352