mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-07-23 11:10:55 -05:00
* server: return 400 instead of 500 on validation error with X-Conversation-Id set_req() attaches the spipe as soon as the header is present, before the request body is parsed. When params validation throws, set_next() never runs and next_orig stays empty, so on_complete() called it and crashed with std::bad_function_call, turning the prepared 400 JSON into a generic 500. on_complete() now treats an empty next_orig as "streaming never started" and evicts the session installed by set_req(), so a failed request leaves nothing behind for discovery or replay. This also covers valid requests that carry the header but do not stream, which previously left an empty finalized session in the map until the GC TTL. * ui: do not send the backend_sampling placeholder On a fresh profile the syncable settings hold the empty string placeholder meaning "let the server decide". Every neighbor field goes through the hasValue() guard that filters it, except backend_sampling, which sent the placeholder verbatim and made every default settings completion fail validation. Guard the field with hasValue() like its neighbors. hasValue(false) is true, so an explicit false still reaches the server and the intent of #18781 (send both true and false) is preserved. Only the placeholder is filtered.
23 KiB
23 KiB