mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-01 02:17:42 -05:00
graph_inputs was populated while splitting the graph, so it only contained the inputs that are used as srcs of some node. With pipeline parallelism (n_copies > 1), each graph input contributes n_copies leafs to graph_copy, so switching between batches that consume different inputs (e.g. token batches that do not use the embeddings input vs image batches that do) changed the graph composition. This shifted the input copies in graph_copy, making the backend ids comparison report spurious changes and forcing the scheduler to re-reserve. The re-reserve could then record smaller input sizes (e.g. out_ids with n_outputs = 0) and abort later on a graph with an unchanged size via GGML_SCHED_DEBUG_REALLOC. Collect the inputs after the split instead, from all input leafs of the graph, so that the graph composition depends only on which inputs exist, not on which inputs are used. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL