Commit Graph

  • 13eeebb1b2 vulkan: subgroup size tuning (llama/12087) Daniele 2025-03-17 12:42:33 +01:00
  • 905b834af1 vulkan: use fp32 in coopmat2 q4_k dequant function (llama/12309) Jeff Bolz 2025-03-17 04:43:35 -05:00
  • 2cd3061a23 vulkan: Pad N dimension of B matrix for coopmat2 perf, to avoid bounds checking (llama/12273) Jeff Bolz 2025-03-17 04:41:59 -05:00
  • 88d59e21b2 vulkan: Adjust coopmat2 tile sizes and selection heuristic (llama/12258) Jeff Bolz 2025-03-17 04:35:00 -05:00
  • 4917f122d4 cmake : enable building llama.cpp using system libggml (llama/12321) Christian Kastner 2025-03-17 10:05:23 +01:00
  • 16a1b77249 SYCL: set extras only on GGML_TYPE_Q4_0 (llama/12366) Akarshan Biswas 2025-03-17 07:15:12 +05:30
  • 51d1398a0a SYCL: Delete redundant plus sign and space (llama/12391) aubreyli 2025-03-15 22:49:03 +08:00
  • 3499dd83c0 SYCL : support non-contiguous tensors in binary ops (add, sub, etc) (llama/12399) fairydreaming 2025-03-15 15:19:30 +01:00
  • 7b7d9ae35e MUL_MAT optimization (llama/12382) Chenguang Li 2025-03-15 09:31:08 +08:00
  • 2dcb7181ff sycl : variable sg_size support for mmvq kernels (llama/12336) Alberto Cabrera Pérez 2025-03-12 09:57:32 +00:00
  • 96ab3b2465 CUDA/HIP: Fix fattn-vec-* when device warp size is not 32 (llama/12315) uvos 2025-03-12 10:14:11 +01:00
  • 08f32992d0 vulkan: fix bug in coopmat1 mul_mat_id (llama/12316) Jeff Bolz 2025-03-12 00:59:19 -05:00
  • 394fae57c3 CUDA/HIP: refractor mmqv to unify the calculation of nwarps and rows per block between host and device code. (llama/12177) uvos 2025-03-11 20:16:03 +01:00
  • 0708835301 ggml-backend : fix backend search path (llama/12330) jklincn 2025-03-11 21:25:17 +08:00
  • 774c519433 metal : Cache the Metal library at the device context level (llama/12265) BB-fat 2025-03-11 19:45:02 +08:00
  • 776cdceb9e mat vec double buffer (llama/12188) Eve 2025-03-10 19:28:11 +00:00
  • 03d050481e musa: support new arch mp_31 and update doc (llama/12296) R0CKSTAR 2025-03-11 01:18:25 +08:00
  • 3d60219622 opencl: use OpenCL C standard supported by the device (llama/12221) Henry Linjamäki 2025-03-10 18:57:00 +02:00
  • 521d72d76e ggml-backend : make path_str compatible with C++20 (llama/12269) Jason C.H 2025-03-09 00:02:39 +08:00
  • 9fb9025a40 ggml : skip intermediate .air file when compiling .metallib (llama/12247) Daniel Bevenius 2025-03-07 14:15:27 +01:00
  • 3c2abb01e8 cmake: Enable specifying exact PowerPC CPU architecture (ggml/1138) Christian Kastner 2025-03-10 19:19:58 +01:00
  • efd9407e22 cmake: Comment out GGML_BIN_DIR for now (ggml/1139) Christian Kastner 2025-03-10 13:06:21 +01:00
  • 3684af2594 scripts : update sync Georgi Gerganov 2025-03-27 10:13:26 +02:00
  • 206459a804 bindings-go : update Makefile to use cmake (#2952) b2280 Daniel Bevenius 2025-03-26 16:21:07 +01:00
  • 21d890d534 whisper : add support for backends with multiple ggml_backend_buffer_type (#2863) b2279 Dan Johansson 2025-03-26 15:54:02 +01:00
  • 0b43a02be8 bindings.java : enable copyLibs task [no ci] (#2949) Daniel Bevenius 2025-03-26 15:01:28 +01:00
  • 2699e1485a bindings.javascript : update test instructions [no ci] (#2951) Daniel Bevenius 2025-03-26 14:49:12 +01:00
  • 594a121f3e readme : add note about SDL2 (#2946) b2276 Page-MS 2025-03-26 03:30:59 -04:00
  • 996581c5e2 whisper.android : add GGML_USE_CPU compile definition (#2945) b2275 Daniel Bevenius 2025-03-25 18:01:18 +01:00
  • 226d344f56 whisper.android.java : update build with ggml source changes (#2942) b2274 Daniel Bevenius 2025-03-25 16:01:59 +01:00
  • bb9f68129f ci: fix SYCL build (#2943) b2273 Akarshan Biswas 2025-03-25 14:50:37 +05:30
  • 30cf30ca82 examples : reduce initial memory to 512MB (#2939) Daniel Bevenius 2025-03-24 14:42:12 +01:00
  • ee6286c35d examples : fix nthread parsing in whisper.wasm (#2938) b2271 Daniel Bevenius 2025-03-24 14:40:00 +01:00
  • c7941d5ccc examples : fix request path for local worker files (#2937) b2270 Daniel Bevenius 2025-03-24 14:33:45 +01:00
  • b82ac32a6c ggml : add logging for native build options/vars (#2935) b2269 Daniel Bevenius 2025-03-24 09:53:38 +01:00
  • edf1ee1ef8 whisper : enhance model download scripts functionality and resolve compiler warning (#2925) b2268 Peter 2025-03-24 19:39:50 +11:00
  • cf5ddb8c21 whisper : initialize decoder's rng with unique seed (#2932) b2267 Daniel Bevenius 2025-03-24 09:36:07 +01:00
  • 7fe4979f25 ci : remove CMAKE_CUDA_ARCHITECTURES in windows-cublas (#2923) b2266 Daniel Bevenius 2025-03-22 15:40:28 +01:00
  • 9bc0dc7235 whisper : update default model download directory behavior to use current working directory when script is in /bin/ directory (#2924) Peter 2025-03-23 01:27:57 +11:00
  • 3fc6ad97a3 whisper.swiftui : Add Core ML support to README [no ci] (#2921) Daniel Bevenius 2025-03-21 11:38:32 +01:00
  • 663cafc1e8 readme : update Python version to 3.11 for Core ML support [no -ci] (#2919) b2263 Daniel Bevenius 2025-03-21 10:31:55 +01:00
  • be9de81171 whisper : add check for CPU backend initialization (#2918) b2262 Daniel Bevenius 2025-03-21 09:53:26 +01:00
  • 21fb513ef1 examples : update whisper.objc README.md (#2916) b2261 Daniel Bevenius 2025-03-21 09:52:53 +01:00
  • 4e56747944 ci : increase windows-cublas evict-old-files to 5d (#2915) b2260 Daniel Bevenius 2025-03-21 08:19:24 +01:00
  • ca75449a92 xcframework : add support for CoreML to ios/macOS (#2912) b2259 Daniel Bevenius 2025-03-20 18:39:08 +01:00
  • 80dad86b2c examples : add WHISPER_SDL2 check to deprecation executables (#2911) b2258 Daniel Bevenius 2025-03-20 18:36:02 +01:00
  • 485ece6725 ci : use ninja and fix caching for windows-cublas (#2910) b2257 Daniel Bevenius 2025-03-20 17:01:48 +01:00
  • e7d9d8687a examples : update wasm examples to include server.py [no ci] (#2908) Daniel Bevenius 2025-03-20 09:07:43 +01:00
  • 6e8242f7fe examples : command.wasm updates (#2904) Daniel Bevenius 2025-03-20 07:02:18 +01:00
  • e27fd6f0c0 ci : refactor cuda toolkit installation steps (#2902) b2254 Daniel Bevenius 2025-03-19 09:41:14 +01:00
  • 96db0c5a9c go : add Encoder Begin Callback (#2900) b2253 Amanda Der Bedrosian 2025-03-19 00:05:04 -07:00
  • d2aaffd5d9 ci : add ccache action to windows-cublas job (#2893) b2252 Daniel Bevenius 2025-03-19 04:53:08 +01:00
  • 215990abde whisper : fix compiler warnings in whisper.cpp (#2895) b2251 Daniel Bevenius 2025-03-18 13:38:41 +01:00
  • 7e23d8c64a ci : add missing env.branch_name to build.yml (#2896) b2250 Daniel Bevenius 2025-03-18 13:38:21 +01:00
  • 740bf7f6a1 whisper : enable compiler warnings for src (#2891) Daniel Bevenius 2025-03-18 05:19:18 +01:00
  • c8e12f59dd ci : add release job and include xcframework (#2889) Daniel Bevenius 2025-03-18 05:18:20 +01:00
  • 83b14c357c examples : use xcframework in whisper.objc example (#2882) Daniel Bevenius 2025-03-17 13:01:24 +01:00
  • 60b481d881 whisper : add option to use system-installed GGML (#2887) Peter 2025-03-17 18:54:48 +11:00
  • 4854789751 convert : update convert-h5-to-ggml.py (#2840) Anders Bjarby 2025-03-17 08:41:05 +01:00
  • e0f3c9d4dd examples : add GGML_USE_CPU=ON flag to whisper.objc (#2880) Daniel Bevenius 2025-03-14 15:40:20 +01:00
  • 1f4886b40d ggml-ci: update input env variables to GG_BUILD_ (#2879) Benjamin Ryan 2025-03-14 03:53:29 -05:00
  • 05ce7476ae ggml-ci: update input env variables to GG_BUILD_ ci/env Benjamin 2025-03-14 03:14:44 -05:00
  • f11de0e73c ggml-ci: add run.sh (#2877) Benjamin Ryan 2025-03-14 02:29:55 -05:00
  • d5cc27ee4d examples : add dl to the list of libraries linked (#2875) Daniel Bevenius 2025-03-14 04:42:20 +01:00
  • 5bb1d58c6a whisper: add xcframework build script (#2873) Martin Destagnol 2025-03-13 02:56:39 -10:00
  • 7d14005717 objc : fix build, tmp remove GPU support, use C++17 Georgi Gerganov 2025-03-08 10:46:37 +02:00
  • 4ffb8e3e4d cmake : fix ggml-config (ggml/0) Georgi Gerganov 2025-03-08 10:33:33 +02:00
  • 1d8d8ae55e sync : ggml Georgi Gerganov 2025-03-08 10:26:00 +02:00
  • eebf6bc0bd ggml-cpu: faster AVX2 variant for IQ1_M (llama/12216) Rémy O 2025-03-07 12:54:22 +01:00
  • dc8f423b40 metal : simplify kernel arguments using a struct (ggml/3229) (llama/12194) BB-fat 2025-03-07 15:35:57 +08:00
  • 548e7052f1 metal : fix default.metallib build (llama/12224) Daniel Bevenius 2025-03-07 06:23:16 +01:00
  • a34cb73dc2 opencl: Noncontiguous norm, rms_norm, disable fp16 for some ops (llama/12217) lhez 2025-03-06 16:20:35 -08:00
  • 82f9496657 cmake : fix undefined reference errors for std::filesystem in ggml (#12092) (llama/12094) xiaofei 2025-03-07 06:58:25 +08:00
  • e3c85e75bd CUDA: fix FA logic for PTX 7.0 and CC >= 7.5 (llama/12222) Johannes Gäßler 2025-03-06 18:45:09 +01:00
  • b9eab73fa2 HIP/CUDA: set the paramerter value in maintain_cuda_graph instead of replaceing it. (llama/12209) uvos 2025-03-06 08:20:52 +01:00
  • 76385c8311 opencl : fix buffer alignment (llama/12197) Henry Linjamäki 2025-03-06 03:33:40 +02:00
  • 442cd1d2e7 opencl : fix ulong kernel args were set from int variables (llama/12174) Henry Linjamäki 2025-03-06 03:31:14 +02:00
  • bc8cb97e02 opencl : fix profile-related errors (llama/12095) simon886212 2025-03-06 09:30:05 +08:00
  • 8dcadf736b ggml-cpu: Faster IQ1 mul_mat_vec on AVX2 using BMI2 instructions (llama/12154) Rémy O 2025-03-06 02:26:10 +01:00
  • 93986b61e0 SYCL: Disable f16 Unary OPs as not supported by the kernels (llama/12201) Akarshan Biswas 2025-03-05 21:28:23 +05:30
  • bd1a9e34c9 ggml : fix GGMLMetalClass ODR (llama/12200) Plamen Minev 2025-03-05 17:16:01 +02:00
  • cc03608e78 ggml : ggml_compute_forward_concat() for arbitrary tensor type (ggml/1118) vmobilis 2025-03-07 11:11:40 +03:00
  • 54a54faee4 vulkan : sync (llama/0) Georgi Gerganov 2025-03-04 21:08:15 +02:00
  • 96a92ecc4c ggml : portability fixes for VS 2017 (llama/12150) mgroeber9110 2025-03-04 17:53:26 +01:00
  • edd1d8686a HIP: implement FlashAttention via rocWMMA for CDNA and RDNA3+ (llama/12032) David Huang 2025-03-04 05:10:54 +08:00
  • dc6f4e7c05 ggml : fix kleidiai build (llama/12159) ag2s20150909 2025-03-03 20:54:08 +08:00
  • 74c85d154e SYCL: Move CPY kernels to a separate file and add few missing kernels (llama/12133) Akarshan Biswas 2025-03-03 15:37:22 +05:30
  • eb2d8b6ffd ggml-backend : keep paths in native string type when possible (llama/12144) Diego Devesa 2025-03-02 22:11:00 +01:00
  • b442dcd598 CUDA: compress mode option and default to size (llama/12029) Erik Scholz 2025-03-01 12:57:22 +01:00
  • c98681e6d5 ggml : upgrade init_tensor API to return a ggml_status (llama/11854) William Tambellini 2025-02-28 05:41:47 -08:00
  • 3bab804981 vulkan: add specific MMV kernels for IQ2 and IQ3 quants + optimizations (llama/11595) Rémy O 2025-02-28 09:42:52 +01:00
  • c927830a70 CUDA: fix logic for V100 + GGML_CUDA_FORCE_MMQ (llama/12098) Johannes Gäßler 2025-02-28 09:26:43 +01:00
  • 992b51b3d5 ggml: aarch64: implement SVE kernels for q2_k_q8_k vector dot (llama/12064) Prashant Vithule 2025-02-28 13:06:12 +05:30
  • 2c882cbe4c CANN: Fix build error with GCC 13 (llama/11990) hipudding 2025-02-28 15:23:47 +08:00
  • 1fbb119b1e vulkan: matmul dequantization improvements (llama/12015) Eve 2025-02-28 07:20:08 +00:00
  • 40dea850fd vulkan: improve im2col (llama/11826) Daniele 2025-02-28 06:52:51 +00:00
  • 8255a830a8 cmake: Fix ggml backend dependencies and installation (llama/11818) Vladimir Vuksanovic 2025-02-27 08:42:48 +01:00
  • a0f76b2da7 vulkan: fix assertion when qy_needs_dequant (llama/12068) Jeff Bolz 2025-02-25 09:30:21 -06:00
  • 394768c48b ggml-cpu: Fix build with sve (llama/12059) Molly Sophia 2025-02-25 19:28:22 +08:00
  • 846e01b2c0 cuda: unary ops as float + de-duplicate (ggml/1130) cmdr2 2025-03-03 20:51:31 +05:30