Files
whisper.cpp/ggml/include
Diego Devesa b11c972b88 llama : separate compute buffer reserve from fattn check (llama/15696)
Exposes ggml_backend_sched_split_graph() to allow splitting the graph without allocating compute buffers and uses it to split the graph for the automatic Flash Attention check.
2025-09-20 13:42:45 +03:00
..
2025-09-20 13:42:39 +03:00