server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)

* server : allow splitting RANK pooling for causal LLM rerankers

Rerank models fall into two categories: bidirectional cross-encoders
(BERT, etc.) that require all tokens in a single physical batch, and
causal LLMs repurposed as rerankers (Qwen3, Qwen3-VL) that can use
chunked prefill like any other decoder.

Previously the server rejected all RANK-pooling inputs larger than
n_ubatch, and the graph builder hardcoded QWEN3/QWEN3VL arch checks to
determine last-token pooling. This broke long-document and multimodal
reranking for causal models.

Fix: expose llama_get_causal_attn(ctx) so the server can check the
effective runtime attention type (reflecting any --attention override
or set_causal_attn call). Also expose llama_model_is_causal(model)
for querying the static architectural property from GGUF metadata.

can_split() now permits chunked prefill for RANK pooling when the
context is causal. The graph builder's inline arch check is replaced
with the same cparams.causal_attn predicate, removing the duplication.

Assisted-by: Opencode/Qwen3.8-27B

* remove unused llama_model_is_causal, fix whitespace

Assisted-by: opencode

---------

Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
This commit is contained in:
Tim Wang
2026-09-27 23:28:10 +02:00
committed by GitHub
co-authored by timothywang21
parent a97cce86a8
commit 4da6337767
5 changed files with 32 additions and 8 deletions
+18 -7
View File
@@ -435,15 +435,26 @@ struct server_slot {
return task->need_embd();
}
// if the context does not have a memory module then all embeddings have to be computed within a single ubatch
// also we cannot split if the pooling would require any past tokens
// (MTP supports splitting — uses task->need_embd() not need_embd())
bool can_split() const {
GGML_ASSERT(task);
return
!task->need_embd() ||
(llama_get_memory(ctx_tgt) && llama_pooling_type(ctx_tgt) == LLAMA_POOLING_TYPE_LAST);
// MTP supports splitting - uses task->need_embd() not need_embd()
if (!task->need_embd()) {
return true;
}
// if the context does not have a memory module then all embeddings have to be computed within a single ubatch
if (!llama_get_memory(ctx_tgt)) {
return false;
}
// context can be chunked/split if the pooling type is LAST
const auto pooling = llama_pooling_type(ctx_tgt);
if (pooling == LLAMA_POOLING_TYPE_LAST) {
return true;
}
// causal rerankers read the last token and have a KV cache, so they can also be chunked/split.
if (pooling == LLAMA_POOLING_TYPE_RANK && llama_get_causal_attn(ctx_tgt)) {
return true;
}
return false;
}
bool can_batch_with(server_slot & other_slot) const {