GGML_RPC_DEBUG is now parsed as a number: 0/unset disables debug logs,
1-3 emit increasingly detailed output (events, per-command trace,
transport detail). Non-numeric values fall back to 1.
The duplicated env/macro blocks in ggml-rpc.cpp and transport.cpp are
replaced by a shared log.h, which becomes the single choke point for
all logging of the RPC backend: LOG_ERROR/LOG_WARN/LOG_INFO for
unconditional severity logs and LOG_DBG/LOG_DBG2/LOG_DBG3 for the
verbosity-gated ones. The transport files no longer need ggml-impl.h,
and the server banner now also goes through the ggml logger (stderr).
Missing logs are added on both the client (handshake, buffer ops,
tensor transfers, graph computes, cache decisions) and the server
(per-command dispatch, graph nodes), including the negotiated
transport via the new socket_t::transport_name().
Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
The original function was broken on Windows for some unicode paths
Paths without a trailing separator now create the last directory too,
matching the function name. All current callers already include a
trailing separator, so this change does not affect them.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* rpc: avoid serializing buffers from other servers
Only include remote buffer pointers when the buffer belongs to the RPC dispatcher receiving the graph. Add a two-server regression test for cross-server tensor serialization.
Assisted-by: Codex
* cont : add ref
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* rpc: support apple RDMA as an RPC transport
* remove set_tensor micro optimization, rpc socket pinning per CR
* remove transparent reconnect
* trigger apple builds on RPC changes
---------
Co-authored-by: Ryan Churaman <rschu@meta.com>
Tests are generally prefixed with -test, so rename export-graph-ops
accordingly.
rpc-server is probably too generic a name for /usr/bin. Because it
should work with any ggml application, it is renamed to ggml-rpc-server.
* Set C locale for consistent float formatting across all binaries.
* Add C locale setting to all tools binaries
Add std::setlocale(LC_NUMERIC, "C") to all 16 binaries in the tools/
directory to ensure consistent floating-point formatting.
* Apply suggestion from @JohannesGaessler
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
* rpc : report actual free memory
Start reporting the free memory on every device instead of using
fixed values. Now llama-cli users can get a nice memory breakdown
when using RPC devices.
* drop --mem in rpc-server
Update the README file to match the newly added functionality of
exposing multiple devices from a single server.
Co-authored-by: Diego Devesa <slarengh@gmail.com>
* rpc : add support for multiple devices
Allow rpc-server to expose multiple devices from a single endpoint.
Change RPC protocol to include device identifier where needed.
closes: #15210
* fixes
* use ggml_backend_reg_t
* address review comments
* fix llama-bench backend report
* address review comments, change device naming
* fix cmd order