mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-08-05 17:40:46 -05:00
* convert text model * main model load ok * convert encoder ok * speaker encoder loading ok * speaker enc graph * adapt vocab for backbone (with some tricks) * add suppress_tokens * poc new mtmd gen api * convert code_predictor to gguf * load gen_code model ok * add clip_encode * wire up * code gen cgraph init version Co-authored-by: Pascal <admin@serveurperso.com> * code2wav convert to gguf * code2wav graph ok * wire up in/out * (wip) subgraph * wire up * wip, correct code2wav * demo (to be removed) * code2wav preserve kv between calls * demo voice clone * llama: add llama_model_get_tok_embd * mtmd_helper_gen_audio API * fix clamp cold prefix Co-authored-by: Pascal <admin@serveurperso.com> * fuse snake op Co-authored-by: Pascal <admin@serveurperso.com> * demo: use proper sampling * update dev docs * polymorphism helper * revamp llama-tts binary * update docs * fix compile * fix lint * nits * add guide + docs * more timings info * clean up code comments * security fixes * update docs * use ggml_build_forward_select, clean up comments * fix ci * use ISO 639-1 language code * rename CODE2WAV --> GEN_WAV, update docs * clean up * clean up tts.cpp * add seq_id * add step_prompt() * mtmd_helper_model_can_chat * clean up comments --------- Co-authored-by: Pascal <admin@serveurperso.com>
5.2 KiB
5.2 KiB
libmtmd dev guide
History
Please refer to multimodal.md for a broader context.
In short:
libmtmdstarted as a wrapper aroundlibllava/clip.cpp- Various components that used to be in
clip.cppare moved progressively to mtmd. For example, preprocessor is now part of mtmd
Terminologies
- mtmd: MulTiMoDal
- bitmap: representing a raw input data, for example: RGB image, PCM audio
- tiles / slices: for llava-uhd-style models, the preprocessor breaks a large input into smaller square images called tiles or slices
- chunk: a mtmd_input_chunk represents a preprocessed input that can then be passed through
mtmd_encode()
Pipeline
A typical pipeline of the core libmtmd is as follows:
- A bitmap (RGB image or PCM audio) is created
- Bitmap and the text prompt is provided to
mtmd_tokenize()that breaks the input into chunks- The tokenizer function first expands a "lazy" bitmap if it finds one. Typically, this is used by video, so that one media token corresponds to one input bitmap
- For models that support "fused" temporal frames like Qwen-VL, the tokenizer tries to merge pair of consecutive frames into one batch
- The preprocessor will then be called, which produces a list of chunks
- Depending on the model itself, special tokens will be injected to separate image chunks (i.e. llava-uhd-style models)
- Multiple bitmaps may be batched together to form a larger
mtmd_batch() - Single image or batch is encoded, via
mtmd_encode()ormtmd_batch_encode() - Get the output embeddings
Helper
We provide a set of helper functions via mtmd_helper to make using libmtmd easier. The helper provides:
- Image, audio and video file decoding (for example, decode raw JPEG into RGB bitmap)
- Manage
llama_batchand calls tollama_decode
Audio generation support
Audio generation is added to mtmd in PR #26254
Currently, we support the 3-stage pipeline below which should cover most TTS models:
- Stage 1: Backbone / Semantic Stage: Backbone model accepts text prompt and reference voice as input
- Stage 2: Acoustic Detail Generator: A model takes the hidden state from backbone and generate audio details (usually as audio codes or mel-spectrogram)
- Stage 3: Waveform Reconstruction: Convert the semantic and acoustic data from previous stages to the final waveform
For example, Qwen3-TTS:
- Reference voice is encoded using ECAPA-TDNN speaker encoder (
speaker_encoder) - Text prompt and reference voice are processed via a backbone (
talker.model) - A model converts sampled semantic token and hidden state from stage 2 into a list of 15 acoustic codes (
talker.code_predictor) - 16 generated codes are converted into waveform (
code2wav)
API design constraints
Due to wide variety of audio generation pipelines, the mtmd_gen_audio system is designed to be flexible and reusable by new models.
mtmd_gen_audio is split into 2 main API:
- Core API
mtmd.h: handles main inference. Important: the API surface must be stateless; caller must handle state management and audio frame accumulation. - Helper API
mtmd-helper.h: provides a model-agnostic stateful API. Usage example can be found in thetools/ttsdirectory.
Checklist for porting new audio generation models to mtmd
- Establish a list of reusable and missing components from the current mtmd implementation.
- For GGUF conversion:
- Backbone model should be converted to a normal text model (loadable via
libllama)- If model used hard-coded embedding row ID, append them to token embeddings and assign token name for them (see
qwen3tts.py) - If model have a specific output logits head for audio codes (usually semantic code), keep the head as-is and pad the logits at inference time (see
src/models/qwen3vl.cpp)
- If model used hard-coded embedding row ID, append them to token embeddings and assign token name for them (see
- Sidecar models (code2wav, bigvgan, etc) must live inside the mmproj GGUF (but can be in different
clip_contextif necessary)- Note: it should use
ggml_build_forward_selectto select graphs if multiple graphs living in the same context
- Note: it should use
- Reuse existing GGUF metadata key name and tensor name whenever possible; think twice before adding extensive changes to GGUF writer. For example, Qwen3-TTS hard-code part of the hparams to
clip.cppas they won't likely to change. - For tensor naming:
- Prefixed with
a.*for tensors used by speaker encoder pipeline - Prefixed with
a.gen.*for generation stages (code / mel-spectrogram / PCM generation)
- Prefixed with
- Backbone model should be converted to a normal text model (loadable via
- Make sure most of the changes happen inside
mtmd-helper-gen.cpp. A good PR looks like this:- 10-20% changes is to add new backbone (text) model and conversion
- 60% changes inside
mtmd-helper-gen.cpp - 10% changes inside
libmtmdandclip.cppsystems - The rest downstream code (CLI, server) should have no changes at all
- Update usage documentation in
tools/tts/README.md
IMPORTANT: If your model needs changes that don't fit the existing infrastructure, open an issue first for discussion.
No-go checklist (these will get the PR rejected and require discussion before proceeding):
- Violating the API design constraints stated above
- Adding a new model-specific binary: the API and binary surface must stay model-agnostic