mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-02 10:57:33 -05:00
llama : preserve original batch order for speculative decoding layer inputs (#29019)
* llama: preserve original batch order for layer inputs Assisted-by: Codex * tests: cover layer-input order across KV layouts Assisted-by: Codex * tests: exercise layer-input ordering on CUDA devices Assisted-by: Codex * llama: make layer input reordering compatible with tensor split Copy each microbatch tensor from offset zero and restore original row order after synchronization. Extend the layer-input regression to cover tensor split and repeated reads and decodes. Assisted-by: Codex * llama: restore token order for unmasked NextN embeddings Use the original-token mapping for unmasked NextN rows, including when layer-input capture is disabled. Keep masked NextN rows on the logits output mapping and preserve offset-zero tensor copies. Extend the existing regression to cover NextN alone, combined layer capture, and masked outputs with repeated decodes and getters. Validation: all 256 CPU/CUDA/tensor configurations pass. Qwen3.8-27B Q4_K_M MTP completes MT-Bench at concurrency 16 before and after. Assisted-by: Codex * ggml: fix WebGPU reservation and OpenVINO hidden-state capture Reserve WebGPU vector attention scratch across batch sizes and refresh reservations when NextN capture settings change. Preserve requested OpenVINO outputs, dynamic shapes, sequence counts, and current graph bindings. Extend existing WebGPU regression coverage and enable strict allocation checks. Assisted-by: Codex * llama: defer regression test and backend fixes to follow-ups Keep this PR focused on restoring token order for layer inputs and unmasked NextN embeddings. Remove the added regression test, OpenVINO and WebGPU changes, and the separate NextN reservation change. Assisted-by: Codex * llama: keep n_embd declaration in its original position Assisted-by: Codex * llama : pass token count to layer input extraction Assisted-by: Codex * llama : name original batch indices batch_idxs Assisted-by: Codex * llama : name extracted embedding indices embd_batch_idxs Assisted-by: Codex * llama : tag target embedding reordering Assisted-by: Codex * llama : tag extraction and name the index capture flag Assisted-by: Codex
This commit is contained in:
@@ -800,6 +800,7 @@ llama_ubatch llama_batch_allocr::ubatch_add(const std::vector<int32_t> & idxs, u
|
||||
udata->seq_idx .resize(LLAMA_MAX_SEQ, -1);
|
||||
udata->output .resize(n_tokens);
|
||||
|
||||
udata->batch_idxs = idxs;
|
||||
udata->seq_id_data.reserve(n_tokens);
|
||||
|
||||
seq_set_t seq_set_unq;
|
||||
|
||||
Reference in New Issue
Block a user