mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-27 08:27:30 -05:00
The reference implementation has two mutually exclusive prompt layouts. In non streaming mode the prefill carries the whole utterance text plus tts_eos summed with codec_pad, and the trailing text hidden collapses to a single tts_pad row. In streaming mode the prefill carries only the first text token and the trailing rows stream the rest of the text followed by tts_eos. The pipeline built the non streaming prefill but the streaming overlay, so the talker saw the utterance a second time during generation and read it twice before emitting codec_eos. The overlay is now the single tts_pad row that matches the prefill.