mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-29 01:17:36 -05:00
Chat goes through protocol adapters, so an OpenAI-compatible endpoint speaks its own wire format: per-backend paths and headers, the model on the request, tools kept on the local server, and token counts synthesized for endpoints that do not stream their own timings. The server store keeps the local server's props while another provider is active, a conversation resolves the provider its model belongs to before sending, and the chat screen never blocks on the local probe when the install has none. Assisted-by: pi:llama.cpp/DeepSeek-V4.1-Flash