Commit Graph

  • 96e1280839 clip : bring back GPU support (#12322) b4870 Xuan-Son Nguyen 2025-03-11 09:20:16 +01:00
  • 2c9f833d17 mat vec double buffer (#12188) b4869 Eve 2025-03-10 19:28:11 +00:00
  • 251364549f musa: support new arch mp_31 and update doc (#12296) b4868 R0CKSTAR 2025-03-11 01:18:25 +08:00
  • 8acdacb3ea opencl: use OpenCL C standard supported by the device (#12221) b4867 Henry Linjamäki 2025-03-10 18:57:00 +02:00
  • 89b2b56e86 readme: added Sidekick to available UIs (#12311) John Bean 2025-03-10 22:13:09 +08:00
  • e128a1bf5b tests : fix test-quantize-fns to init the CPU backend (#12306) b4865 Georgi Gerganov 2025-03-10 14:07:15 +02:00
  • 6ef79a67ca common : refactor '-o' option (#12278) b4864 marcoStocchi 2025-03-10 12:34:13 +01:00
  • 4e39a3c332 server: extract <think> tags from qwq outputs (#12297) b4863 Olivier Chafik 2025-03-10 10:59:03 +00:00
  • be421fc429 tool-call: ensure there's always a non-empty tool call id (#12292) Olivier Chafik 2025-03-10 09:45:29 +00:00
  • 87c2630546 allow missing content in message if tool_calls provided (#12293) b4861 Olivier Chafik 2025-03-10 09:45:07 +00:00
  • 2b3a25c212 sampler: fixes trigger tokens + lazy grammars (fix typo cast from token to string) (#12291) b4860 Olivier Chafik 2025-03-10 09:44:42 +00:00
  • 8352cdc87b llava : fix bug in minicpm-v code (#11513) b4859 tc-mb 2025-03-10 16:33:24 +08:00
  • 1e2f78a004 server : add speculative decoding presets for FIM (#12287) Georgi Gerganov 2025-03-09 19:08:20 +02:00
  • 87dae2fd15 Vulkan: Print coopmat shapes, then exit 0cc4m/vulkan-print-coopmat-shapes 0cc4m 2025-03-09 10:53:55 +00:00
  • 0fd7ca7a21 authors : update (#12271) Georgi Gerganov 2025-03-08 18:26:00 +02:00
  • 6fefc05a7a ggml-backend : make path_str compatible with C++20 (#12269) b4856 Jason C.H 2025-03-09 00:02:39 +08:00
  • 25840747e6 Vulkan: Add device architecture enum and logic to recognize AMD generations 0cc4m/vulkan-device-architecture 0cc4m 2025-03-08 08:04:45 +00:00
  • 7ab364390f server : infill gen ends on new line (#12254) b4855 Georgi Gerganov 2025-03-07 20:54:30 +02:00
  • f27c1afc40 ggml-quants : improve TQ2_0 imatrix Francis Couture-Harpin 2025-03-07 12:54:56 -05:00
  • c75753a01b server : infill gen ends on new line gg/server-infill-end-on-nl Georgi Gerganov 2025-03-07 17:19:55 +02:00
  • 7c7f3b7f43 ggml : skip intermediate .air file when compiling .metallib (#12247) b4854 Daniel Bevenius 2025-03-07 14:15:27 +01:00
  • 102ac1891d sync : ggml b4853 Georgi Gerganov 2025-03-07 14:00:27 +02:00
  • d6ae2fa061 ggml : ggml_compute_forward_concat() for arbitrary tensor type (ggml/1118) vmobilis 2025-03-07 11:11:40 +03:00
  • 68d0027f3d ggml-cpu: faster AVX2 variant for IQ1_M (#12216) b4851 Rémy O 2025-03-07 12:54:22 +01:00
  • ea002810a2 ci : fix save-load test invocations (#12245) Georgi Gerganov 2025-03-07 12:19:31 +02:00
  • aefa65e442 ci : fix save-load test invokations gg/ci-fix-save-load Georgi Gerganov 2025-03-07 12:17:33 +02:00
  • 8fad3c7a7c server : Log original chat template parsing error (#12233) b4849 Sigbjørn Skjæret 2025-03-07 11:15:33 +01:00
  • aae2903e0b clang-tidy : disable bugprone-branch-clone gg/clang-tidy-disable-bugprone Georgi Gerganov 2025-03-07 11:36:55 +02:00
  • 7cf64f6bee sync: minja - support QwQ-32B (#12235) b4848 Olivier Chafik 2025-03-07 09:33:37 +00:00
  • 5e2d57b2b2 metal : simplify kernel arguments using a struct (#3229) (#12194) b4847 BB-fat 2025-03-07 15:35:57 +08:00
  • f1648e91cf HIP: fix rocWMMA build flags under Windows (#12230) b4846 David Huang 2025-03-07 15:06:08 +08:00
  • d6c95b0740 metal : fix default.metallib build (#12224) Daniel Bevenius 2025-03-07 06:23:16 +01:00
  • d76a86d967 opencl: Noncontiguous norm, rms_norm, disable fp16 for some ops (#12217) lhez 2025-03-06 16:20:35 -08:00
  • 776f9e59cc cmake : fix undefined reference errors for std::filesystem in ggml (#12092) (#12094) xiaofei 2025-03-07 06:58:25 +08:00
  • 3d652bfddf readme : update bindings (#12229) Lucas Moura Belo 2025-03-06 16:15:13 -03:00
  • 5220a16d18 CUDA: fix FA logic for PTX 7.0 and CC >= 7.5 (#12222) Johannes Gäßler 2025-03-06 18:45:09 +01:00
  • 3ffbbd5ce1 HIP: rocWMMA documentation and enabling in workflow builds (#12179) David Huang 2025-03-06 21:14:11 +08:00
  • 42994048a3 update function-calling.md w/ template override for functionary-small-v3.2 (#12214) Olivier Chafik 2025-03-06 09:03:31 +00:00
  • e9b2f84f14 llava: add big-endian conversion for image encoder (#12218) Aaron Teo 2025-03-06 16:33:21 +08:00
  • e721c05c93 HIP/CUDA: set the paramerter value in maintain_cuda_graph instead of replaceing it. (#12209) b4837 uvos 2025-03-06 08:20:52 +01:00
  • 57b6abf85a android : fix KV cache log message condition (#12212) b4836 Han Yin 2025-03-05 22:22:49 -08:00
  • 94bb63e4f0 opencl : fix buffer alignment (#12197) b4835 Henry Linjamäki 2025-03-06 03:33:40 +02:00
  • f79243992c opencl : fix ulong kernel args were set from int variables (#12174) b4834 Henry Linjamäki 2025-03-06 03:31:14 +02:00
  • ed4ce0dda2 opencl : fix profile-related errors (#12095) b4833 simon886212 2025-03-06 09:30:05 +08:00
  • 07d1572347 ggml-cpu: Faster IQ1 mul_mat_vec on AVX2 using BMI2 instructions (#12154) b4832 Rémy O 2025-03-06 02:26:10 +01:00
  • 5e43f104cc SYCL: Disable f16 Unary OPs as not supported by the kernels (#12201) b4831 Akarshan Biswas 2025-03-05 21:28:23 +05:30
  • 16e4b22c5e ggml : fix GGMLMetalClass ODR (#12200) b4830 Plamen Minev 2025-03-05 17:16:01 +02:00
  • 074c4fd39d ci : add fetch-depth to xcframework upload (#12195) b4829 Daniel Bevenius 2025-03-05 14:16:40 +01:00
  • 669912d9a5 tool-call: fix Qwen 2.5 Coder support, add micro benchmarks, support trigger patterns for lazy grammars (#12034) Olivier Chafik 2025-03-05 13:05:13 +00:00
  • fa31c438e0 ci : fix xcframework artifact tag (#12191) b4827 Daniel Bevenius 2025-03-05 10:22:29 +01:00
  • 3ccbfe5a71 ci : remove xframework upload (#12190) b4826 Daniel Bevenius 2025-03-05 08:34:02 +01:00
  • 06a92a193a server : fix cache reuse logic (#12161) Clauszy 2025-03-05 15:25:45 +08:00
  • a057897ad4 llama : add xcframework build script (#11996) b4824 Daniel Bevenius 2025-03-05 06:30:31 +01:00
  • 5bbe6a9fe9 ggml : portability fixes for VS 2017 (#12150) b4823 mgroeber9110 2025-03-04 17:53:26 +01:00
  • 20a9b8f5e1 readme : fix roadmap link (#12185) Georgi Gerganov 2025-03-04 18:42:44 +02:00
  • 56d7a9f812 main: allow preloading conversation with -p and add -st / --single-turn (#12145) b4821 Sigbjørn Skjæret 2025-03-04 17:19:39 +01:00
  • 1a24c4621f server: fix deadly typo in response_format.json_schema.schema handling (#12168) b4820 Olivier Chafik 2025-03-04 06:24:07 +00:00
  • becade5de7 HIP: implement FlashAttention via rocWMMA for CDNA and RDNA3+ (#12032) b4819 David Huang 2025-03-04 05:10:54 +08:00
  • dfd6b2c0be sync : ggml b4818 Georgi Gerganov 2025-03-03 17:57:38 +02:00
  • b64d7cc272 cuda: unary ops as float + de-duplicate (ggml/1130) cmdr2 2025-03-03 20:51:31 +05:30
  • 3d1cf3cf33 sync : ggml Georgi Gerganov 2025-02-28 12:37:35 +02:00
  • 0cbee131ad cuda/vulkan: specify fp32-only support for some operations in supports_op (ggml/1129) cmdr2 2025-02-28 12:36:46 +02:00
  • 8371d44595 sync : ggml Georgi Gerganov 2025-02-28 09:09:58 +02:00
  • 87abb7e903 cuda/cpu: Increase support for fp16 unary operations (ggml/1125) cmdr2 2025-02-28 12:34:39 +05:30
  • 6d4c23b81b whisper : support GGML_BACKEND_DL (whisper/2843) Diego Devesa 2025-02-27 13:35:07 +01:00
  • 6512a90037 cmake : fix compile assumptions for power9/etc (whisper/2777) midnight 2025-02-05 04:41:10 -08:00
  • 4512055792 Told cmake to install ggml-cpp.h as a public header file. (ggml/1126) petterreinholdtsen 2025-02-26 21:44:00 +01:00
  • f54a4ba11e Support pure float16 add/sub/mul/div operations in the CUDA (and CPU) backend (ggml/1121) cmdr2 2025-02-25 18:06:34 +05:30
  • aede2074f6 scripts : sync-ggml-am.sh fix Georgi Gerganov 2025-02-28 09:09:38 +02:00
  • 2679c3b55d ci : set GITHUB_ACTION env var for server tests (#12162) Daniel Bevenius 2025-03-03 16:17:36 +01:00
  • c43af9276b tts: add speaker file support (#12048) b4806 dm4 2025-03-03 21:09:29 +08:00
  • d5c63cd7f9 test-backend-ops : add option -p to filter by op params (#12155) b4805 Diego Devesa 2025-03-03 14:00:46 +01:00
  • 9660ffef58 ggml : fix kleidiai build (#12159) b4804 ag2s20150909 2025-03-03 20:54:08 +08:00
  • c950a1f692 Adding UTF-8 support to llama.cpp (#12111) b4803 Eric Curtin 2025-03-03 12:44:56 +00:00
  • 7b69003af7 webui : add ?m=... and ?q=... params (#12148) Xuan-Son Nguyen 2025-03-03 11:42:45 +01:00
  • ece9745bb8 SYCL: Move CPY kernels to a separate file and add few missing kernels (#12133) b4801 Akarshan Biswas 2025-03-03 15:37:22 +05:30
  • cc473cac7c ggml-backend : keep paths in native string type when possible (#12144) b4800 Diego Devesa 2025-03-02 22:11:00 +01:00
  • 14dec0c2f2 main: use jinja chat template system prompt by default (#12118) b4799 Sigbjørn Skjæret 2025-03-02 14:53:48 +01:00
  • 46596caf6d apply various in places Xuan Son Nguyen 2025-03-01 20:42:18 +01:00
  • 1d6ba97789 remove token_info API Xuan Son Nguyen 2025-03-01 16:21:16 +01:00
  • 1782cdfed6 main: update outdated system prompt message (followup to #12131) (#12132) b4798 Sigbjørn Skjæret 2025-03-01 15:22:27 +01:00
  • 1170135dfb llama_batch_ext_add_text Xuan Son Nguyen 2025-03-01 14:00:14 +01:00
  • 40989f4116 correct llama_decode_ext Xuan Son Nguyen 2025-03-01 14:00:05 +01:00
  • 45a8e76745 common : add --system-prompt parameter, replace behavior of -p in conversation mode (#12131) b4797 Sigbjørn Skjæret 2025-03-01 13:56:45 +01:00
  • 80c41ddd8f CUDA: compress mode option and default to size (#12029) b4796 Erik Scholz 2025-03-01 12:57:22 +01:00
  • 9e75c49d35 Merge branch 'master' into xsn/private_batch_api Xuan Son Nguyen 2025-03-01 12:13:03 +01:00
  • f0ffd81130 adapt common Xuan Son Nguyen 2025-03-01 12:12:52 +01:00
  • 2cc4a5e44a webui : minor typo fixes (#12116) Vivian 2025-03-01 15:45:09 +05:30
  • 624f7bd03b graph : add comments gg/llama-kv-cache Georgi Gerganov 2025-02-28 21:13:08 +02:00
  • 0f7daa9d1b graph : move non-context related logic to llm_build_context Georgi Gerganov 2025-02-28 19:56:10 +02:00
  • 06c2b1561d convert : fix Norway problem when parsing YAML (#12114) Xuan-Son Nguyen 2025-02-28 17:44:46 +01:00
  • 9cab53c7dd cont : migrate the rest of the inputs out of llama_context Georgi Gerganov 2025-02-28 18:01:25 +02:00
  • 7f02ee562e context : decouple inputs, llama_graph_i become const (WIP) Georgi Gerganov 2025-02-28 14:09:20 +02:00
  • 70680c48e5 ggml : upgrade init_tensor API to return a ggml_status (#11854) b4793 William Tambellini 2025-02-28 05:41:47 -08:00
  • c43a3e7996 llama : add Phi-4-mini support (supersede #12099) (#12108) b4792 Xuan-Son Nguyen 2025-02-28 12:44:11 +01:00
  • 84d5f4bc19 Update granite vision docs for 3.2 model (#12105) Alex Brooks 2025-02-28 04:31:47 -07:00
  • 38db8a5861 llama : introduce concept of llama_memory Georgi Gerganov 2025-02-28 10:51:17 +02:00
  • 438a83926a vulkan: add specific MMV kernels for IQ2 and IQ3 quants + optimizations (#11595) b4790 Rémy O 2025-02-28 09:42:52 +01:00
  • 9c42b1718c CUDA: fix logic for V100 + GGML_CUDA_FORCE_MMQ (#12098) b4789 Johannes Gäßler 2025-02-28 09:26:43 +01:00
  • 05e6f5aad0 ggml: aarch64: implement SVE kernels for q2_k_q8_k vector dot (#12064) b4788 Prashant Vithule 2025-02-28 13:06:12 +05:30