mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-07-28 05:30:58 -05:00
Reuse the buffer for the ggml context which is used for creating the compute graph on the server side. This partially addresses a memory leak created by the CUDA backend due to using buffer addresses as cache keys. ref: #21265 ref: #20315