mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-10-03 03:17:32 -05:00
* Make the drafter probabilistic and the target verify by rejection sampling * Drop stale spec_draft_q before drafting * Fallback to argmax sampling for grammar-constrained requests and adding flag for enabling probabilistic draft sampling. Default flag value is greedy. * Support grammar-constrained requests in rejection sampling * Fix - renormalize distribution after masking * copy rng on sampler copy and re-accept drafted tokens on replay * Fix draft sampler sharing the target's rng stream * Simplify the rejection sampler's inputs and move replay to the server * Truncate the draft candidates along with the draft --------- Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com> Co-authored-by: Pranesh Gonegandla <pgonegandla@nvidia.com>