Commit Graph

  • 7538246e7c cuda : add f32 to bf16 copy op (#12806) b5083 Sigbjørn Skjæret 2025-04-08 23:21:31 +02:00
  • 06e1d3119a convert : write tensors in parallel Francis Couture-Harpin 2025-04-08 16:31:45 -04:00
  • b32efad2bc llava: improve clip_ctx destructor to not memleak load_image_size (#12834) b5082 Matt Clayton 2025-04-08 16:01:58 -04:00
  • a19b5cef16 llama : fix FA when KV cache is not used (i.e. embeddings) (#12825) b5081 Georgi Gerganov 2025-04-08 19:54:51 +03:00
  • 78a1ba0a4f server : fix thread.join() on exit (#12831) b5080 Xuan-Son Nguyen 2025-04-08 18:37:06 +02:00
  • 2dabf759e7 llava: add more helper functions to check projector types in clip context (#12824) b5079 dm4 2025-04-08 21:49:13 +08:00
  • 1d343b4069 arg : Including limits file on AIX (#12822) b5078 Prajwal B Mehendarkar 2025-04-08 18:00:59 +05:30
  • 8ca6e1c3a4 server : webui : Improve Chat Input with Auto-Sizing Textarea (#12785) characharm 2025-04-08 14:14:59 +05:00
  • 656babd6c2 Revert "sycl:remove redundant memcopy in function ggml_backend_sycl_buffer_set_tensor" (#12812) b5076 Neo Zhang Jianyu 2025-04-08 15:03:21 +08:00
  • a226bc7a9a gguf-py : support lazy tensor splitting (#12809) compilade 2025-04-08 03:03:07 -04:00
  • e9e1882d2d rm tail space revert-12734-fix_code_in_ggmlsycl Neo Zhang Jianyu 2025-04-08 13:43:11 +08:00
  • 76f2ed3d77 Update ggml/src/ggml-sycl/ggml-sycl.cpp Neo Zhang Jianyu 2025-04-08 13:16:14 +08:00
  • d271172ab1 Update ggml/src/ggml-sycl/ggml-sycl.cpp Neo Zhang Jianyu 2025-04-08 10:32:18 +08:00
  • 564a05daf2 Revert "sycl: remove redundant memcopy in function ggml_backend_sycl_buffer_s…" Neo Zhang Jianyu 2025-04-08 10:29:41 +08:00
  • da140da72a gguf-py : fix flake8 lint compilade/lazy-tuples Francis Couture-Harpin 2025-04-07 19:38:35 -04:00
  • 6cbbd8e1df gguf-py : support lazy tensor splitting Francis Couture-Harpin 2025-04-07 19:20:54 -04:00
  • 1466621e73 llama : Support llama 4 text-only (#12791) b5074 Xuan-Son Nguyen 2025-04-07 23:06:44 +02:00
  • 82974011f3 opencl: better identify Adreno GPU (#12760) b5073 lhez 2025-04-07 13:22:54 -07:00
  • 4ccea213bc hellaswag: display estimated score confidence interval (#12797) b5072 stduhpf 2025-04-07 17:47:08 +02:00
  • 1a1ab7e7a4 cuda : fix HIP and MUSA BF16 (#0) b5071 Georgi Gerganov 2025-04-07 13:18:07 +03:00
  • a4e46e28f9 sync : ggml Georgi Gerganov 2025-04-07 12:32:39 +03:00
  • ff067dbcb9 ggml : simplify Arm fp16 CPU logic (ggml/1177) Georgi Gerganov 2025-04-07 12:25:15 +03:00
  • 36ca8b3628 CUDA: don't convert BF16 weights to FP32 (ggml/1174) Sigbjørn Skjæret 2025-04-04 21:05:12 +02:00
  • 995083e4ed cpu: move all the operators into a separate c++ file (except mul_mat) (ggml/1167) cmdr2 2025-04-02 17:46:16 +05:30
  • 518a01480e sycl: remove redundant memcopy in function ggml_backend_sycl_buffer_set_tensor (#12734) b5066 zhouwg 2025-04-07 23:22:57 +08:00
  • e391d3ee8d ci : no curl on ggml-ci (#12796) Xuan-Son Nguyen 2025-04-07 14:37:28 +02:00
  • ced26486ff cont sync-ggml-25-04-03-try-fix Georgi Gerganov 2025-04-07 15:24:01 +03:00
  • bd3f59f812 cmake : enable curl by default (#12761) b5064 Xuan-Son Nguyen 2025-04-07 13:35:19 +02:00
  • 52b3d71f12 CANN: fix typo in ggml-cann (#12733) zhouwg 2025-04-07 19:34:14 +08:00
  • 5ef588ba58 test Georgi Gerganov 2025-04-07 13:18:07 +03:00
  • 6232ceec72 sync : ggml Georgi Gerganov 2025-04-07 12:32:39 +03:00
  • e638450acd ggml : simplify Arm fp16 CPU logic (ggml/1177) Georgi Gerganov 2025-04-07 12:25:15 +03:00
  • 4683cb402a CUDA: don't convert BF16 weights to FP32 (ggml/1174) Sigbjørn Skjæret 2025-04-04 21:05:12 +02:00
  • 53cb49e337 cpu: move all the operators into a separate c++ file (except mul_mat) (ggml/1167) cmdr2 2025-04-02 17:46:16 +05:30
  • d0d5b2232b CANN: Refactor to reduce duplicate code (#12731) b5062 hipudding 2025-04-07 17:10:36 +08:00
  • 916c83bfe7 musa: fix compilation warnings in mp_22/31 (#12780) b5061 R0CKSTAR 2025-04-06 21:23:54 +08:00
  • 0c74b04376 vulkan: fix NaN issue in flash attention shader (#12776) b5060 Jeff Bolz 2025-04-06 04:03:47 -05:00
  • 80b717d493 vulkan: Use unclamped loads for flash attention mask (#12720) b5059 Jeff Bolz 2025-04-06 03:47:13 -05:00
  • 6bf28f0111 Vulkan: Tune Vulkan mmq int dot shader for performance (#12767) b5058 0cc4m 2025-04-05 18:04:03 +02:00
  • f1e3eb4249 common : fix includes in arg.cpp and gemma3-cli.cpp (#12766) b5057 Sergey Fedorov 2025-04-05 23:46:00 +08:00
  • 0364178ca2 clip : refactor clip_init, add tests (#12757) b5056 Xuan-Son Nguyen 2025-04-05 17:17:40 +02:00
  • c6ff5d2a8d common: custom hf endpoint support (#12769) b5055 エシュナヴァリシア 2025-04-05 21:31:42 +08:00
  • 7a84777f42 sync: minja (#12739) b5054 Olivier Chafik 2025-04-04 13:16:39 -07:00
  • 3e1d29348b kv-cache : simplify + fix warning for recurrent models (#12756) b5053 Georgi Gerganov 2025-04-04 21:48:10 +03:00
  • 1be76e4620 ci: add Linux cross-compile build (#12428) b5052 bandoti 2025-04-04 14:05:12 -03:00
  • b772394297 server : webui : Upgrade daisyui, tailwindcss. (#12735) Nauful Shaikh 2025-04-04 09:09:52 -05:00
  • 23106f94ea gguf-split : --merge now respects --dry-run option (#12681) b5050 nick huang 2025-04-04 22:09:12 +08:00
  • 94148ba330 sycl: allow ggml-sycl configuration and compilation using Visual Studio project/solution (#12625) b5049 Nicolò Scipione 2025-04-04 16:00:46 +02:00
  • 9ac4d611d0 cmake: fix ggml-shaders-gen compiler paths containing spaces (#12747) Ronny Brendel 2025-04-04 15:12:40 +02:00
  • fe564b0dfb ci : rename job MSVC -> MinGW gg/ci-rename-job Georgi Gerganov 2025-04-04 13:51:48 +03:00
  • 43ab09b85d ci : testing (wip) gg/ci-add-arm-msvc-toolchain Georgi Gerganov 2025-04-04 13:43:43 +03:00
  • 7a73e861a7 cont gg/arm-try-fix-msvc Georgi Gerganov 2025-04-04 12:00:23 +03:00
  • 1b07edfb56 ggml : trying stuff (wip) Georgi Gerganov 2025-04-04 11:33:44 +03:00
  • 348888e0dc docs : add XCFramework section to README.md [no ci] (#12746) Daniel Bevenius 2025-04-04 10:24:12 +02:00
  • 74d4f5b041 vulkan: Hybrid waitForFences/getFenceStatus to reduce fence latency (#12630) b5046 Jeff Bolz 2025-04-04 00:54:35 -05:00
  • 35e592eb30 vulkan: set cmake minimum and project name in vulkan-shaders (#12744) b5045 Jeff Bolz 2025-04-04 00:53:20 -05:00
  • 7d7b1bafa7 opencl: update doc for OpenCL (#12702) lhez 2025-04-03 22:18:17 -07:00
  • c262beddf2 CUDA: Prefer vector flash decoding kernel for Gemma models (#12738) b5043 Gaurav Garg 2025-04-03 21:50:29 +05:30
  • 5dd5d1ab00 vocab : use string_view::find() to avoid unnecessary looking up beyond the fragment range (#12706) yumeyao 2025-04-03 23:32:54 +08:00
  • 1c059995e0 vulkan: Fix missing cmake logic for dot product extension (#12721) b5041 Jeff Bolz 2025-04-03 10:08:26 -05:00
  • 2004644b7a ci : add env variable in ggml-ci and document the same in SYCL.md (#12736) Atharva Dubey 2025-04-03 13:12:39 +01:00
  • 5f696e88e0 sync : minja (inclusionAI/Ling) and update tests (#12699) b5039 R0CKSTAR 2025-04-03 19:51:35 +08:00
  • 819b7d7cce sync : ggml Georgi Gerganov 2025-04-03 10:49:44 +03:00
  • 0c8dad10a3 cpu: move all the operators into a separate c++ file (except mul_mat) (ggml/1167) cmdr2 2025-04-02 17:46:16 +05:30
  • 193c3e03a6 fix MUSA compiler warning (#12704) b5038 a3sh 2025-04-03 15:32:55 +08:00
  • 65cfe136a0 CANN: Support operator SIN COS ARGMAX (#12709) b5037 Chenguang Li 2025-04-03 15:18:08 +08:00
  • 3f9da22c2b Simplify and improve CUDA graphs through use of indirect copy pointers (#9017) b5036 Alan Gray 2025-04-03 02:31:15 +01:00
  • 2a0dc97e56 CANN: Fix failed test cases (#12708) b5035 hipudding 2025-04-03 08:49:51 +08:00
  • 97a20c012b opencl: use max_alloc_size in backend ctx instead of querying again (#12705) b5034 lhez 2025-04-02 17:01:42 -07:00
  • f01bd02376 vulkan: Implement split_k for coopmat2 flash attention. (#12627) b5033 Jeff Bolz 2025-04-02 14:25:08 -05:00
  • 6f3bd38640 cmake: remove caching from vulkan coopmat checks (#12719) b5032 bandoti 2025-04-02 14:56:26 -03:00
  • be0a0f8cae vulkan: Implement grouped query attention in the coopmat2 FA shader (#12559) b5031 Jeff Bolz 2025-04-02 12:40:32 -05:00
  • 92e3006bb6 Vulkan: Fix mmq int dot float cache size (#12722) b5030 0cc4m 2025-04-02 19:12:30 +02:00
  • 833e2b7409 model : print tensor size during load (#12711) b5029 Georgi Gerganov 2025-04-02 16:38:54 +03:00
  • e0e912f49b llama : add option to override model tensor buffers (#11397) b5028 Diego Devesa 2025-04-02 14:52:01 +02:00
  • a10b36c91a llama : refactor kv cache guard (#12695) Georgi Gerganov 2025-04-02 14:32:59 +03:00
  • 83a88bd6af vocab : BailingMoE : change possessive quantifiers to greedy (#12677) b5026 Sigbjørn Skjæret 2025-04-02 11:21:48 +02:00
  • 42eb248f46 common : remove json.hpp from common.cpp (#12697) b5025 Xuan-Son Nguyen 2025-04-02 09:58:34 +02:00
  • 9bacd6b374 [CANN] get_rows and dup optimization (#12671) Chenguang Li 2025-04-02 15:22:13 +08:00
  • 267c1399f1 common : refactor downloading system, handle mmproj with -hf option (#12694) Xuan-Son Nguyen 2025-04-01 23:44:05 +02:00
  • f423981ac8 opencl : fix memory allocation size (#12649) b5022 Junil Kim 2025-04-02 01:54:34 +09:00
  • e39e727e9a llama : use LLM_KV_GENERAL_FILE_TYPE instead of gguf_find_key (#12672) b5021 jklincn 2025-04-01 20:54:28 +08:00
  • 5936a616e4 convert : BailingMoE : fix qkv split when head_dim is 0 (#12687) Sigbjørn Skjæret 2025-04-01 14:37:13 +02:00
  • 3fd072a540 metal : use F32 prec in FA kernels (#12688) b5019 Georgi Gerganov 2025-04-01 14:57:19 +03:00
  • a6f32f0b34 Fix clang warning in gguf_check_reserved_keys (#12686) b5018 R0CKSTAR 2025-04-01 19:12:53 +08:00
  • 2bb3597e42 vulkan: fix build when glslc doesn't support coopmat (#12683) b5017 Wagner Bruna 2025-04-01 06:38:07 -03:00
  • 8293970542 SYCL: Rename oneMKL to oneMath (#12192) b5016 Romain Biessy 2025-04-01 10:24:29 +02:00
  • 8bbf26083d SYCL: switch to SYCL namespace (#12674) b5015 Akarshan Biswas 2025-04-01 13:41:39 +05:30
  • 35782aeedb convert : BailingMoE : avoid setting rope_dim to 0 (#12678) Sigbjørn Skjæret 2025-03-31 23:09:48 +02:00
  • c80a7759da vocab : add special infill tokens for CodeLlama (#11850) b5013 Daniel Bevenius 2025-03-31 18:40:56 +02:00
  • 250d7953e8 ggml : faster ssm scan (#10558) b5012 a3sh 2025-04-01 00:05:13 +08:00
  • 403fbacbbc convert : Qwerky : use lora_rank_tokenshift and lora_rank_decay if present (#12667) Sigbjørn Skjæret 2025-03-31 16:36:25 +02:00
  • a8a1f33567 Vulkan: Add DP4A MMQ and Q8_1 quantization shader (#12135) b5010 0cc4m 2025-03-31 14:37:01 +02:00
  • 1790e73157 cmake : fix whitespace (#0) b5009 Georgi Gerganov 2025-03-31 15:05:30 +03:00
  • 0114a32da0 sync : ggml Georgi Gerganov 2025-03-31 14:59:21 +03:00
  • a7724480fd cmake: improve Vulkan cooperative matrix support checks (whisper/2966) Sandro Hanea 2025-03-31 12:44:36 +02:00
  • 1a85949067 llava : proper description fix (#12668) b5006 Sigbjørn Skjæret 2025-03-31 11:28:30 +02:00
  • 6c02a032fa SYCL: Remove misleading ggml_sycl_op_flatten function (#12387) b5005 Akarshan Biswas 2025-03-31 14:55:24 +05:30
  • f52d59d771 llava : fix clip loading GGUFs with missing description (#12660) b5004 Sigbjørn Skjæret 2025-03-31 11:07:07 +02:00
  • 52de2e5949 tts : remove printfs (#12640) b5003 marcoStocchi 2025-03-31 10:20:30 +02:00