mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-07-26 12:40:58 -05:00
* Use F16 for memory_k and memory_v * add command line switch to use f16 instead of f32 for memory k+v --------- Co-authored-by: Ty Everett <ty@tyweb.us>
20 KiB
20 KiB