mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-07-29 14:11:13 -05:00
With -hfd pointing to a repo that ships mtp-/dflash-/eagle3- sidecars and no --spec-type given, the draft resolved to a full model while the sidecar was the intended draft. When the speculative types are still at their default, discover the sidecars of the draft repo, pick the first available following the existing mtp > dflash > eagle3 priority, and set the corresponding type, so this now works without any extra flag: llama-server -hf repo:Q3_K_M -hfd repo:Q8_0 An explicit --spec-type disables the inference, and a draft repo without sidecars keeps resolving to a full model as before.
203 KiB
203 KiB