Commit Graph

  • 5218ea21b8 cuda : fix dmmv cols requirement to 2*GGML_CUDA_DMMV_X (llama/8800) slaren 2024-08-01 15:26:22 +02:00
  • e60be821ce added android implementation of ggml_print_backtrace_symbols (llama/8751) l3utterfly 2024-07-30 23:40:18 +09:00
  • 19708df884 cann: update cmake (llama/8765) wangshuai09 2024-07-30 18:37:35 +08:00
  • 3f190addda Add TIMESTEP_EMBEDDING OP (llama/8707) zhentaoyu 2024-07-30 14:56:51 +08:00
  • b355ee7cfa ggml: bugfix: fix the inactive elements is agnostic for risc-v vector (llama/8748) CarterLi999 2024-07-30 00:38:34 +08:00
  • 49ac8872b4 cuda : organize vendor-specific headers into vendors directory (llama/8746) R0CKSTAR 2024-07-29 20:56:12 +08:00
  • 8ef98ae7e3 add conv support (llama/8688) Meng, Hengyu 2024-07-29 10:50:27 +08:00
  • e471adcfa5 feat: Support Moore Threads GPU (llama/8383) R0CKSTAR 2024-07-28 07:41:25 +08:00
  • aa816c922c ggml : ignore more msvc warnings (ggml/906) Borislav Stanimirov 2024-08-07 10:00:56 +03:00
  • b3264eb266 metal : fix struct name (ggml/912) Georgi Gerganov 2024-08-07 09:57:00 +03:00
  • eb2eb87a58 metal : add abort callback (ggml/905) Conrad Kramer 2024-08-07 02:55:49 -04:00
  • 83fcb0e486 vulkan : implement Stable Diffusion operators (ggml/904) 0cc4m 2024-08-04 17:28:08 +02:00
  • f7bb412878 ggml : move c parameter comment to ggml_rope_ext (ggml/901) Daniel Bevenius 2024-07-29 15:06:06 +02:00
  • ef6dcf0d0c ggml : resolve sync conflicst (ggml/0) Georgi Gerganov 2024-07-27 17:17:23 +03:00
  • c7ea4fd235 common : handle new quant types (ggml/0) Georgi Gerganov 2024-07-27 17:17:04 +03:00
  • 525f190917 ggml : add ggml-aarch64 (ggml/0) Dibakar Gope 2024-07-27 17:16:40 +03:00
  • dd916a2852 ggml : reduce hash table reset cost (llama/8698) slaren 2024-07-27 04:41:55 +02:00
  • 0620fe00ec ggml: handle ggml_init failure to fix NULL pointer deref (llama/8692) DavidKorczynski 2024-07-25 22:23:05 +01:00
  • 31d0a9a14f fix multi-gpu issue on sycl (llama/8554) Chen Xi 2024-07-25 11:45:18 +00:00
  • c06970dd72 ggml : add and use ggml_cpu_has_llamafile() (llama/8664) Georgi Gerganov 2024-07-25 12:37:42 +03:00
  • 7598acf525 Re-add erroneously removed -fsycl from GGML_EXTRA_LIBS (llama/8667) Joe Todd 2024-07-24 11:55:26 +01:00
  • 43ddfce969 sycl : Add support for non-release DPC++ & oneMKL (llama/8644) Joe Todd 2024-07-23 14:58:37 +01:00
  • a7e6d2cd9c Vulkan IQ4_NL Support (llama/8613) 0cc4m 2024-07-23 10:56:49 +02:00
  • 86506b0c5c Allow all RDNA2 archs to use sdot4 intrinsic (llama/8629) Jeroen Mostert 2024-07-23 10:50:40 +02:00
  • 11182fae34 fix scratch size of softmax (llama/8642) luoyu-intel 2024-07-23 07:43:28 +00:00
  • 0bc8bffe1d ggml: fix compile error for RISC-V (llama/8623) Mark Zhuang 2024-07-22 15:56:45 +08:00
  • 8c4f30497a CUDA: MMQ code deduplication + iquant support (llama/8495) Johannes Gäßler 2024-07-20 22:25:26 +02:00
  • b1ee3a8444 gguf : handle null name during init (llama/8587) Georgi Gerganov 2024-07-20 17:15:42 +03:00
  • be9a16fd3f ggml : fix quant dot product with odd number of blocks (llama/8549) slaren 2024-07-19 17:17:27 +02:00
  • f4d9a95b0f ggml : add friendlier error message to fopen errors (llama/8575) Clint Herron 2024-07-19 07:05:45 -04:00
  • a8ab3abe09 CUDA: fix partial offloading for ne0 % 256 != 0 (llama/8572) Johannes Gäßler 2024-07-18 23:48:47 +02:00
  • fb6a835938 cmake : install all ggml public headers (llama/8480) 65a 2024-07-18 07:47:12 -07:00
  • 8923bb4292 Add Ascend NPU backend (llama/6035) hipudding 2024-07-17 19:23:50 +08:00
  • fcba6aa352 make/cmake: add missing force MMQ/cuBLAS for HIP (llama/8515) Johannes Gäßler 2024-07-16 21:20:59 +02:00
  • 8807fe608b Refactor lora adapter support (llama/8332) Xuan Son Nguyen 2024-07-15 20:50:47 +02:00
  • 3e94c7a81d add concat through dim 1/2 (llama/8483) Meng, Hengyu 2024-07-15 19:32:15 +08:00
  • 77af3254e1 Vulkan MMQ Fix (llama/8479) 0cc4m 2024-07-15 09:38:52 +02:00
  • d4b3cffec4 vulkan : cmake integration (llama/8119) bandoti 2024-07-13 13:12:39 -03:00
  • b852a4c5ca metal : template-ify some of the kernels (llama/8447) Georgi Gerganov 2024-07-13 18:32:33 +03:00
  • 2157abaab4 ggml : minor naming changes (llama/8433) Georgi Gerganov 2024-07-12 10:46:02 +03:00
  • 68d609a12c fix the mul_mat_id ut issues (llama/8427) Chen Xi 2024-07-12 00:52:04 +00:00
  • 5a8ae474f0 ggml : add NVPL BLAS support (ggml/8329) (llama/8425) Nicholai Tukanov 2024-07-11 11:49:15 -05:00
  • 84493d7f3e cuda : suppress 'noreturn' warn in no_device_code (llama/8414) Daniel Bevenius 2024-07-11 17:53:42 +02:00
  • 15d71189e9 CUDA: optimize and refactor MMQ (llama/8416) Johannes Gäßler 2024-07-11 16:47:47 +02:00
  • 37e962580f Use multi_ptr to clean up deprecated warnings (llama/8256) AidanBeltonS 2024-07-10 16:10:49 +01:00
  • db0ea7a2f2 ggml : move sgemm sources to llamafile subfolder (llama/8394) Georgi Gerganov 2024-07-10 15:23:29 +03:00
  • 5498b0e6c0 ggml : add AArch64 optimized GEMV and GEMM Q4 kernels (llama/5780) Dibakar Gope 2024-07-10 07:14:51 -05:00
  • 2af4a52c39 sycl : Reenabled mmvq path for the SYCL Nvidia Backend (llama/8372) Alberto Cabrera Pérez 2024-07-09 15:03:15 +01:00
  • eee2fe882e sycl : fix powf call in device code (llama/8368) Alberto Cabrera Pérez 2024-07-08 14:22:41 +01:00
  • 0d1a11e5e2 ggml : loop tiling optimizations for scalar path (ggml/898) Mahesh Madhav 2024-07-25 00:54:08 -07:00
  • b2ead7d6f4 ggml: add support for float16 input tensors in pooling operations (ggml/895) Ivan Filipov 2024-07-22 14:32:02 +03:00
  • 8da6fd4dff vulkan : initialize vk_buffer_struct members to VK_NULL_HANDLE (ggml/893) Tony Wasserka 2024-07-20 20:49:44 +02:00
  • ab8ec9e940 cmake : only enable GGML_NATIVE and x86 flags if not crosscompiling (ggml/885) Borislav Stanimirov 2024-07-12 17:24:20 +03:00
  • 701265bf38 scripts : sync new files (#0) Georgi Gerganov 2024-08-08 14:00:51 +03:00
  • fe36c90971 cmake : fix compile in xcode (#2311) Daven Sanassy 2024-08-05 07:48:26 +01:00
  • 6739eb83c3 whisper : handle empty mel (#2324) Georgi Gerganov 2024-07-27 20:35:04 +03:00
  • f68298ce06 whisper : use vulkan as gpu backend when available (#2302) Matt Stephenson 2024-07-16 03:21:09 -04:00
  • 7ae885c1ef whisper : fix DTW assert (#2299) arizhih 2024-07-15 14:50:36 +02:00
  • d207c68822 cmake : use WHISPER_EXTRA_FLAGS (#2294) Georgi Gerganov 2024-07-09 18:54:18 +03:00
  • 16d72504fe cmake : allow external ggml Borislav Stanimirov 2024-07-08 17:08:55 +03:00
  • 1c31f9d4a8 cmake : try to fix openvino build (#2281) Georgi Gerganov 2024-07-08 15:36:51 +03:00
  • 8ecb2f1f68 cmake : remove install of llama convert script [no ci] (#2266) Georgi Gerganov 2024-07-08 14:21:04 +03:00
  • 5226c3d45c make : remove llama prints [no ci] (#2265) Georgi Gerganov 2024-07-08 14:19:36 +03:00
  • dbf9c15e30 talk-llama : sync llama.cpp Georgi Gerganov 2024-07-08 14:14:17 +03:00
  • d3f6c34976 examples : fix compile warnings [no ci] (#0) Georgi Gerganov 2024-07-08 14:09:09 +03:00
  • 425e2910a3 sync : ggml Georgi Gerganov 2024-07-08 13:50:28 +03:00
  • 49868aa851 ggml : sync sycl (skip) (#0) Georgi Gerganov 2024-07-08 13:50:14 +03:00
  • ff08e30ab5 scripts : fix sync scripts Georgi Gerganov 2024-07-08 13:48:14 +03:00
  • 95f2a191c0 ggml : remove unnecessary UNUSED macro call (ggml/880) Daniel Bevenius 2024-07-08 12:03:42 +02:00
  • 00422ec3cf cmake : add GGML_BUILD and GGML_SHARED macro definitions (llama/8281) Natsu 2024-07-05 22:29:35 +08:00
  • c5b05321e9 Enabled more data types for oneMKL gemm_batch (llama/8236) Ouadie EL FAROUKI 2024-07-05 13:23:25 +01:00
  • 5dc636a65a CUDA: MMQ support for iq4_nl, iq4_xs (llama/8278) Johannes Gäßler 2024-07-05 09:06:31 +02:00
  • 73703a144f CUDA: revert part of the RDNA1 optimizations (llama/8309) Daniele 2024-07-05 07:06:09 +00:00
  • e89fdceec2 CUDA: fix MMQ stream-k rounding if ne00 % 128 != 0 (llama/8311) Johannes Gäßler 2024-07-05 09:05:34 +02:00
  • 29a2739d27 Fix WARP_SIZE=16 bug of Intel GPU (llama/8266) luoyu-intel 2024-07-05 05:06:13 +00:00
  • ee6d17f6b4 rm get_work_group_size() by local cache for performance (llama/8286) Neo Zhang Jianyu 2024-07-05 10:32:29 +08:00
  • 95e90823d9 Define and optimize RDNA1 (llama/8085) Daniele 2024-07-03 23:02:58 +00:00
  • 005cc45df3 fix typo (llama/8267) Judd 2024-07-03 20:40:16 +08:00
  • c2c60dc9ba Removes multiple newlines at the end of files that is breaking the editorconfig step of CI. (llama/8258) Clint Herron 2024-07-02 12:18:10 -04:00
  • 4af3194b7c cuda : update supports_op for matrix multiplication (llama/8245) slaren 2024-07-02 08:39:38 +02:00
  • 4a2ba1a065 Fix win build conflict of math library (llama/8230) luoyu-intel 2024-07-02 04:50:07 +00:00
  • f096cc6807 Fix the sub group size of Intel (llama/8106) luoyu-intel 2024-07-02 02:16:00 +00:00
  • e4bc83ab47 CUDA: refactor and optimize IQ MMVQ (llama/8215) Johannes Gäßler 2024-07-01 20:39:06 +02:00
  • db7e0dbe6e Update SYCL-Rope op and Refactor (llama/8157) zhentaoyu 2024-07-01 19:39:06 +08:00
  • bf88c94da9 CUDA: fix MMQ stream-k for --split-mode row (llama/8167) Johannes Gäßler 2024-06-27 16:26:05 +02:00
  • 3eea171cab feat: cuda implementation for ggml_conv_transpose_1d (ggml/854) John Balis 2024-07-02 11:09:52 -05:00
  • 64a56ebf13 ci : disable java build Georgi Gerganov 2024-07-08 14:26:59 +03:00
  • bec9836849 server : add inference path to make OAI API compatible (#2270) Emmanuel Schmidbauer 2024-07-08 07:24:58 -04:00
  • c118733a29 sync : ggml + fix sync script Georgi Gerganov 2024-06-26 23:20:19 +03:00
  • bb3dd45524 make : disable CUDA graphs Georgi Gerganov 2024-06-26 23:20:13 +03:00
  • 04e7fa6f4f ggml : add GGML_CUDA_USE_GRAPHS option, restore GGML_CUDA_FORCE_CUBLAS (cmake) (llama/8140) slaren 2024-06-26 21:34:14 +02:00
  • 9f7f36d4c9 make : disable CUDA mel build Georgi Gerganov 2024-06-26 22:25:25 +03:00
  • 4a62efbb95 cmake : minor fixes Georgi Gerganov 2024-06-26 21:42:39 +03:00
  • 0a55a70b9b make : fix missing -O3 Georgi Gerganov 2024-06-26 21:20:45 +03:00
  • ceb77363cd ggml : disable CUDA graphs for non-llama.cpp projects gg/disable-cuda-graphs Georgi Gerganov 2024-06-26 20:14:22 +03:00
  • dc8cc2dd6f whisper : disable CUDA mel + fix FFMPEG Georgi Gerganov 2024-06-26 20:11:38 +03:00
  • 3efedb9511 sync : ggml Georgi Gerganov 2024-06-26 19:40:23 +03:00
  • e30c679928 whisper : reorganize source code + improve CMake (#2256) Georgi Gerganov 2024-06-26 19:34:09 +03:00
  • bf4cb4abad whisper : optimize fft() function (#2242) mky_coder 2024-06-18 23:10:33 +08:00
  • e293f17d34 talk-llama : sync llama.cpp Georgi Gerganov 2024-06-18 09:45:37 +03:00