GGML_RPC_DEBUG is now parsed as a number: 0/unset disables debug logs, 1-3 emit increasingly detailed output (events, per-command trace, transport detail). Non-numeric values fall back to 1. The duplicated env/macro blocks in ggml-rpc.cpp and transport.cpp are replaced by a shared log.h, which becomes the single choke point for all logging of the RPC backend: LOG_ERROR/LOG_WARN/LOG_INFO for unconditional severity logs and LOG_DBG/LOG_DBG2/LOG_DBG3 for the verbosity-gated ones. The transport files no longer need ggml-impl.h, and the server banner now also goes through the ggml logger (stderr). Missing logs are added on both the client (handshake, buffer ops, tensor transfers, graph computes, cache decisions) and the server (per-command dispatch, graph nodes), including the negotiated transport via the new socket_t::transport_name(). Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
Overview
Important
This example and the RPC backend are currently in a proof-of-concept development stage. As such, the functionality is fragile and insecure. Never run the RPC server on an open network or in a sensitive environment!
The ggml-rpc-server allows exposing ggml devices on a remote host.
The RPC backend communicates with one or several instances of ggml-rpc-server and offloads computations to them.
This can be used for distributed LLM inference with llama.cpp in the following way:
flowchart TD
rpcb<-->|TCP|srva
rpcb<-->|TCP|srvb
rpcb<-.->|TCP|srvn
subgraph hostn[Host N]
srvn[ggml-rpc-server]<-.->dev4["CUDA0"]
srvn[ggml-rpc-server]<-.->dev5["CPU"]
end
subgraph hostb[Host B]
srvb[ggml-rpc-server]<-->dev3["Metal"]
end
subgraph hosta[Host A]
srva[ggml-rpc-server]<-->dev["CUDA0"]
srva[ggml-rpc-server]<-->dev2["CUDA1"]
end
subgraph host[Main Host]
local["Local devices"]<-->ggml[llama-cli]
ggml[llama-cli]<-->rpcb[RPC backend]
end
style hostn stroke:#66,stroke-width:2px,stroke-dasharray: 5 5
classDef devcls fill:#5B9BD5
class local,dev,dev2,dev3,dev4,dev5 devcls
By default, ggml-rpc-server exposes all available accelerator devices on the host.
If there are no accelerators, it exposes a single CPU device.
Usage
Remote hosts
On each remote host, build the backends for each accelerator by adding -DGGML_RPC=ON to the build options.
For example, to build the ggml-rpc-server with support for CUDA accelerators:
mkdir build-rpc-cuda
cd build-rpc-cuda
cmake .. -DGGML_CUDA=ON -DGGML_RPC=ON
cmake --build . --config Release
When started, the ggml-rpc-server will detect and expose all available CUDA devices:
$ bin/ggml-rpc-server
ggml_cuda_init: GGML_CUDA_FORCE_MMQ: no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 1 CUDA devices:
Device 0: NVIDIA GeForce RTX 5090, compute capability 12.0, VMM: yes
Starting RPC server v3.0.0
endpoint : 127.0.0.1:50052
local cache : n/a
Devices:
CUDA0: NVIDIA GeForce RTX 5090 (32109 MiB, 31588 MiB free)
You can control the set of exposed CUDA devices with the CUDA_VISIBLE_DEVICES environment variable or the --device command line option. The following two commands have the same effect:
$ CUDA_VISIBLE_DEVICES=0 bin/ggml-rpc-server -p 50052
$ bin/ggml-rpc-server --device CUDA0 -p 50052
Main host
On the main host build llama.cpp with the backends for the local devices and add -DGGML_RPC=ON to the build options.
Finally, when running llama-cli or llama-server, use the --rpc option to specify the host and port of each ggml-rpc-server:
$ llama-cli -hf ggml-org/gemma-3-1b-it-GGUF -ngl 99 --rpc 192.168.88.10:50052,192.168.88.11:50052
By default, llama.cpp distributes model weights and the KV cache across all available devices -- both local and remote -- in proportion to each device's available memory.
You can override this behavior with the --tensor-split option and set custom proportions when splitting tensor data across devices.
Local cache
The RPC server can use a local cache to store large tensors and avoid transferring them over the network.
This can speed up model loading significantly, especially when using large models.
To enable the cache, use the -c option:
$ bin/ggml-rpc-server -c
By default, the cache is stored in the $HOME/.cache/llama.cpp/rpc directory and can be controlled via the LLAMA_CACHE environment variable.
RDMA transport
The RPC backend can use RDMA instead of TCP for lower latency and higher throughput. The transport is negotiated during the initial handshake -- no changes to command-line usage are required, and the connection falls back to TCP unless both peers can use RDMA.
Two providers are supported, each enabled by default when its library is found at build time:
- Linux: RoCEv2-capable NICs (e.g. Mellanox ConnectX), via
libibverbs. - macOS: RDMA over Thunderbolt on Apple silicon Macs with Thunderbolt 5, via
librdma. Requires macOS 26.2 or later, with RDMA enabled once from macOS Recovery viardma_ctl enable. See TN3205.
RDMA is point-to-point, so each side uses the local device whose GID matches the address the connection was made on. Connect over the RDMA-capable link -- with Thunderbolt, use the peer's Thunderbolt address in --rpc; a connection made over another interface stays on TCP.
To force plain TCP without rebuilding, set GGML_RPC_NO_RDMA on either peer:
$ GGML_RPC_NO_RDMA=1 bin/ggml-rpc-server
Troubleshooting
The GGML_RPC_DEBUG environment variable controls the verbosity of the logs emitted by the RPC backend.
It can be set on the server, on the client (e.g. llama-cli), or both. Larger values produce more detailed output:
- unset /
0- disabled (only warnings and errors are printed) 1- high-level events: connections, handshake, buffer operations, tensor transfers, graph computes2- per-command trace: every RPC message sent/received, cache hits/misses, queue events3- transport detail: byte counts, per-command timings, graph node details
Non-numeric values are treated as 1.
$ GGML_RPC_DEBUG=1 bin/ggml-rpc-server
$ GGML_RPC_DEBUG=2 bin/llama-cli -hf ggml-org/gemma-3-1b-it-GGUF -ngl 99 --rpc 192.168.88.10:50052