mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-29 09:27:33 -05:00
* context : do not re-reserve the scheduler when toggling causal_attn `llama_context::set_causal_attn()` marks the scheduler to do a full re-reserve on every change of the flag. For vision inputs, this flag is flipped twice around each non-causal image chunk for Gemma models, resulting in two expensive `sched_reserve()` passes per image. This is especially slow for multi-image or video inputs. The cost of a re-reserve scales with context and ubatch configurations, so larger settings pay more per image (see table below). The re-reserve is unnecessary in this case because `causal_attn` only changes the values written to KQ mask, not tensor shapes or any other buffer sizes. Note: `causal_attn` is a graph reuse key (`llm_graph_params` via `cparams`), so a new graph is built regardless of `sched_need_reserve`, so this doesn't change the graph rebuilding behaviour. llama-server with gemma-4-26B-A4B Q4_0 + BF16 mmproj, 130-token images, cache_prompt=false, prompt_ms median of 3 (before -> after): | images | config | H200 before -> after | RTX 4090 before -> after | |-|-|-|-| | 1 | `-c 8192 -ub 512` | 134 -> 105 ms (1.27×) | 201 -> 119 ms (1.69×) | | 24 | `-c 8192 -ub 512` | 2278 -> 1562 ms (1.46×) | 3559 -> 1748 ms (2.04×) | | 24 | `-c 32768 -ub 2048` | 5379 -> 1584 ms (3.40×) | 13377 -> 1759 ms (7.61×) | Generated output remains identical before and after. * qwen4exp : make the indexer bias shape independent of causal_attn The block/cell bias path was selected on cparams.causal_attn, so the causal and non-causal graphs differed in tensor shapes and ops. With the re-reserve removed (previous commit), a runtime flip resulted in reallocating the compute buffers, which would fail under GGML_SCHED_NO_REALLOC. This commit selects the block path from the mask shape only, independent of causal_attn. causal_attn is instead passed to set_input_qsa. causal_attn is fixed per graph as it's part of the reuse key. Causal values are unchanged. Non-causal values now follow the reference rule, where every visible block competes on score and only unpooled cells are always selected. * context : state the causal_attn shape rule in the comment * cont : add TODOs --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>