Commit Graph

  • fc49ee4479 ruby : support new-segment callback (#2506) KITAITI Makoto 2024-10-28 22:43:27 +09:00
  • c0ea41f6b2 ruby : add Metal support (#2516) KITAITI Makoto 2024-10-28 20:08:09 +09:00
  • 0fbaac9c89 whisper : fix index overflow in token-level timestamp logic (#2505) Josscii 2024-10-23 20:14:03 +08:00
  • a5abfe6a90 readme : update links and make commands (#2489) toboil-features 2024-10-17 13:25:18 +03:00
  • d3f7137cc9 ruby : fix bindings (#2484) KITAITI Makoto 2024-10-17 00:44:04 +09:00
  • f7c99e49b3 readme : add Vulkan notice (#2488) toboil-features 2024-10-16 18:43:26 +03:00
  • 1d5752fa42 make : fix GGML_VULKAN=1 build (#2485) Georgi Gerganov 2024-10-16 18:42:47 +03:00
  • b6049060dd whisper : add dtw preset for large-v3-turbo (#2481) Rotem Dan 2024-10-15 21:00:21 +03:00
  • 06a1da9daf convert : handle max_target_positions (#2477) CrispStrobe 2024-10-14 09:46:33 +02:00
  • 746d173592 readme : update the Quick Start section (#2475) Salman Faroz 2024-10-14 13:14:57 +05:30
  • fdbfb460ed whisper : add OpenVINO init with state (#2464) Sandro Hanea 2024-10-08 19:08:00 +02:00
  • ebca09a3d1 release : v1.7.1 v1.7.1 Georgi Gerganov 2024-10-07 13:06:48 +03:00
  • 9f346d0084 vulkan : retry allocation with fallback flags (#2451) SRHMorris 2024-10-06 08:34:20 +01:00
  • 6a94163b91 release : v1.7.0 v1.7.0 Georgi Gerganov 2024-10-05 16:43:26 +03:00
  • 8a35b58c4f scripts : bench v3-turbo Georgi Gerganov 2024-10-05 16:22:53 +03:00
  • 1789abca84 whisper : remove mel leftover constants (396089f) Georgi Gerganov 2024-10-05 16:13:03 +03:00
  • 847f94fdeb whisper : zero-out the KV cache upon clear (#2445) Georgi Gerganov 2024-10-05 15:22:17 +03:00
  • 6e40108a59 objc : fix build Georgi Gerganov 2024-10-05 15:18:50 +03:00
  • 1ba185f4af metal : zero-init buffer contexts (#0) Georgi Gerganov 2024-10-05 14:33:54 +03:00
  • 396089f3cf whisper : revert mel-related changes (#0) Georgi Gerganov 2024-10-05 14:29:45 +03:00
  • 941912467d whisper : adapt to latest ggml (skip) (#0) Georgi Gerganov 2024-10-05 13:14:03 +03:00
  • 0b1b094a67 ggml : fix typo in example usage ggml_gallocr_new (ggml/984) Daniel Bevenius 2024-10-04 15:46:18 +02:00
  • 40e52a76b9 ggml : fixes after sync (ggml/983) Diego Devesa 2024-10-04 08:41:40 +02:00
  • cf977670e6 ggml-backend : add device and backend reg interfaces (llama/9707) Diego Devesa 2024-10-03 21:25:11 +03:00
  • df2c364de7 Fixed dequant precision issues in Q4_1 and Q5_1 (llama/9711) Ouadie EL FAROUKI 2024-10-03 07:50:44 +01:00
  • 1acfadb721 ggml-backend : add device and backend reg interfaces (llama/9707) Diego Devesa 2024-10-03 01:49:47 +02:00
  • ea642144d2 Initial cmake support of SYCL for AMD GPUs (llama/9658) Alberto Cabrera Pérez 2024-10-02 13:57:18 +01:00
  • 282a8654c4 vulkan : do not use tensor->extra (llama/9407) Radoslav Gerganov 2024-10-02 13:49:16 +03:00
  • 936cf3beb7 ggml/ex: calculate accuracy in graph, adapt MNIST (ggml/980) Johannes Gäßler 2024-10-03 17:29:59 +02:00
  • bc92c2f8f0 ggml: refactor cross entropy loss CPU impl. (ggml/976) Johannes Gäßler 2024-10-02 15:32:39 +02:00
  • f7d55e0614 scripts : sync ggml-backend.cpp Georgi Gerganov 2024-10-05 13:09:36 +03:00
  • f62a546e03 whisper : fix excessive memory usage (#2443) Georgi Gerganov 2024-10-05 12:36:40 +03:00
  • 2944cb72d9 examples : update dr_wav.h to newer version (#2449) Rahul Vadhyar 2024-10-04 13:34:51 +05:30
  • ccc2547210 talk-llama : sync llama.cpp Georgi Gerganov 2024-10-02 15:14:46 +03:00
  • 162a455402 metal : reduce command encoding overhead (llama/9698) Georgi Gerganov 2024-10-02 15:12:16 +03:00
  • ff2cb0811f sync : ggml Georgi Gerganov 2024-10-02 15:11:43 +03:00
  • 5e9d6baa48 test: fix OPT_STEP_ADAMW for test-backend-ops (ggml/974) Johannes Gäßler 2024-09-30 09:55:23 +02:00
  • 845f8d663e vulkan : mul_mat: fix UB with small warps (ggml/952) Salvatore Mesoraca 2024-09-30 09:14:09 +02:00
  • 31fdf05fda ggml : fix ggml_cast (ggml/973) Borislav Stanimirov 2024-09-30 10:11:41 +03:00
  • 0ac6666cd2 ggml: fix gradient allocation logic (ggml/966) Johannes Gäßler 2024-09-29 23:18:02 +02:00
  • 6c91da80b8 ggml : define missing HWCAP flags (llama/9684) Georgi Gerganov 2024-09-29 21:18:23 +03:00
  • c245168ba3 ggml : add run-time detection of neon, i8mm and sve (llama/9331) Dan Johansson 2024-09-28 14:06:16 +02:00
  • 280fee8fa0 Enable use to the rebar feature to upload buffers to the device. (llama/9251) Markus Tavenrath 2024-09-28 12:05:05 +02:00
  • 78b4c1c25f mtgpu: enable VMM (llama/9597) R0CKSTAR 2024-09-26 09:27:40 +08:00
  • 1edea2eb4b ggml : remove assert for AArch64 GEMV and GEMM Q4 kernels (llama/9217) Charles Xu 2024-09-25 15:12:20 +02:00
  • 96808786b7 cann: fix crash when llama-bench is running on multiple cann devices (llama/9627) Dou Xinpeng 2024-09-25 11:30:38 +08:00
  • bb57ecb85e CUDA: remove bad assert (ggml/972) Johannes Gäßler 2024-09-29 19:56:17 +02:00
  • abdb73c7cc vulkan : multithread pipeline creation (ggml/963) Jeff Bolz 2024-09-29 11:50:17 -05:00
  • 391e548a43 vulkan : fix build for GGML_VULKAN_RUN_TESTS, add TFLOPS to log (ggml/961) Jeff Bolz 2024-09-27 02:58:01 -05:00
  • 2a29afd4c6 vulkan : argsort barriers must be under uniform control flow (ggml/951) Salvatore Mesoraca 2024-09-26 08:59:42 +02:00
  • 5963004ff9 ggml : fix GGML_MAX_N_THREADS + improve formatting (ggml/969) Georgi Gerganov 2024-09-24 13:23:59 +03:00
  • ede1718f6d server : ffmpeg overwrite leftover temp file (#2431) gilbertgong 2024-10-02 05:06:40 -07:00
  • 2ef717b293 whisper : add large-v3-turbo (#2440) Georgi Gerganov 2024-10-01 15:57:06 +03:00
  • 8feb375fbd tests : remove test-backend-ops (#2434) Georgi Gerganov 2024-09-27 11:48:33 +03:00
  • 69339af2d1 ci : disable failing CUDA and Java builds Georgi Gerganov 2024-09-25 10:03:34 +03:00
  • 0d2e2aed80 readme : fix references to download-ggml-model.sh (#2427) Hugo 2024-09-24 20:07:51 +02:00
  • 451e9ee92c make : remove "talk" target until updated Georgi Gerganov 2024-09-24 14:15:09 +03:00
  • 1133ac98a8 ggml : add ggml-cpu-impl.h (skip) (#0) Georgi Gerganov 2024-09-24 13:27:33 +03:00
  • 76d27eec9a sync : ggml Georgi Gerganov 2024-09-24 13:23:04 +03:00
  • fe18c29ab8 talk-llama : sync llama.cpp Georgi Gerganov 2024-09-24 13:22:55 +03:00
  • 234f9bd320 ggml : add AVX512DQ requirement for AVX512 builds (llama/9622) Eric Zhang 2024-09-24 16:03:21 +08:00
  • 3b183cfae7 log : add CONT level for continuing previous log entry (llama/9610) Georgi Gerganov 2024-09-24 10:15:35 +03:00
  • 02285dff81 threads: fix msvc build without openmp (llama/9615) Max Krasnyansky 2024-09-23 21:18:48 -07:00
  • 2fc1d20f9e cuda: add q8_0->f32 cpy operation (llama/9571) Ivan 2024-09-24 03:14:24 +03:00
  • 08e8414f27 threads: improve ggml_barrier scaling with large number of threads (llama/9598) Max Krasnyansky 2024-09-23 11:42:43 -07:00
  • 05c6139625 ggml : AVX512 gemm for Q4_0_8_8 (llama/9532) Srihari-mcw 2024-09-23 19:36:38 +05:30
  • 896c41ef30 metal : use F32 prec for K*Q in vec FA (llama/9595) Georgi Gerganov 2024-09-23 11:27:47 +03:00
  • c36ddc43c6 Revert "[SYCL] fallback mmvq (ggml/9088)" (llama/9579) Akarshan Biswas 2024-09-23 08:58:06 +05:30
  • 13f41af43e musa: enable building fat binaries, enable unified memory, and disable Flash Attention on QY1 (MTT S80) (llama/9526) R0CKSTAR 2024-09-22 22:55:49 +08:00
  • 3fc5306b82 Fix merge error in #9454 (llama/9589) Molly Sophia 2024-09-22 21:26:50 +08:00
  • adf2474b10 CUDA: enable Gemma FA for HIP/Pascal (llama/9581) Johannes Gäßler 2024-09-22 09:34:52 +02:00
  • 008816a257 RWKV v6: RWKV_WKV op CUDA implementation (llama/9454) Molly Sophia 2024-09-22 10:29:12 +08:00
  • 33e5a6612e ggml-alloc : fix list of allocated tensors with GGML_ALLOCATOR_DEBUG (llama/9573) slaren 2024-09-21 14:24:23 +02:00
  • f0a7d65b3d Update CUDA graph on scale change plus clear nodes/params (llama/9550) agray3 2024-09-21 01:41:07 +01:00
  • 54e5095765 examples : adapt to ggml.h changes (ggml/0) Georgi Gerganov 2024-09-20 21:50:16 +03:00
  • 34291099fb ggml : refactoring (llama/#0) Georgi Gerganov 2024-09-20 21:24:06 +03:00
  • d245d7aec7 ggml : fix builds (llama/0) Georgi Gerganov 2024-09-20 20:12:52 +03:00
  • d661283e68 ggml : fix trailing whitespace (llama/0) Georgi Gerganov 2024-09-20 19:13:02 +03:00
  • c0761c95f5 CUDA: fix sum.cu compilation for CUDA < 11.7 (llama/9562) Johannes Gäßler 2024-09-20 18:35:35 +02:00
  • 138e20b697 ggml : fix n_threads_cur initialization with one thread (llama/9538) slaren 2024-09-18 19:13:08 +02:00
  • a8d9abfa22 threadpool : skip polling for unused threads (llama/9461) Max Krasnyansky 2024-09-17 01:19:46 -07:00
  • 195afd6dc1 ggml : link MATH_LIBRARY not by its full path (llama/9339) Michael Podvitskiy 2024-09-16 13:06:50 +02:00
  • 1fd78999e8 cmake : do not hide GGML options + rename option (llama/9465) Georgi Gerganov 2024-09-16 10:27:50 +03:00
  • 374e9e0c5e ggml : IQ4_NL sgemm + Q4_0 AVX optimization (llama/9422) Eve 2024-09-16 06:48:24 +00:00
  • a2cb5b4183 metal : handle zero-sized allocs (llama/9466) Georgi Gerganov 2024-09-16 09:05:56 +03:00
  • 288ae5176e common : reimplement logging (llama/9418) Georgi Gerganov 2024-09-15 20:46:12 +03:00
  • d868122a5a cmake : correct order of sycl flags (llama/9497) Michael Podvitskiy 2024-09-15 18:55:52 +02:00
  • 2ba25fb122 cmake : try to fix sycl+intel build (llama/9487) Michael Podvitskiy 2024-09-15 09:06:38 +02:00
  • 4f4687cb74 ggml : ggml_type_name return "NONE" for invalid values (llama/9458) Yuri Khrustalev 2024-09-14 05:54:37 -04:00
  • 66b00fad0d cmake : use list(APPEND ...) instead of set() + dedup linker (llama/9463) Georgi Gerganov 2024-09-14 10:55:05 +03:00
  • c6cc8d16c3 cann: Add host buffer type for Ascend NPU (llama/9406) Dou Xinpeng 2024-09-12 19:46:43 +08:00
  • 3f8f8a78a2 riscv : modify Makefile and add a RISCV_VECT to print log info (llama/9442) Ahmad Tameem 2024-09-12 16:24:31 +05:00
  • 3e47686919 cann: Fix error when running a non-exist op (llama/9424) Xinpeng Dou 2024-09-12 09:02:35 +08:00
  • a53b69a003 CUDA: fix --split-mode row race condition (llama/9413) Johannes Gäßler 2024-09-11 10:22:40 +02:00
  • d1c9b47360 musa: remove Clang builtins mapping (llama/9421) R0CKSTAR 2024-09-11 09:46:55 +08:00
  • 32f659861a sycl : update support conditions (llama/9394) Alberto Cabrera Pérez 2024-09-11 01:53:42 +01:00
  • a785232bf9 metal : fix compile warning with GGML_METAL_NDEBUG (llama/0) Georgi Gerganov 2024-09-10 10:17:03 +03:00
  • 0677293503 rpc : fix segfault with nkvo (llama/9389) Radoslav Gerganov 2024-09-09 18:40:10 +03:00
  • 1fbdb813c0 ggml : vector length agnostic SVE support (llama/9290) Prashant Vithule 2024-09-09 21:07:18 +05:30
  • 67725ac8f3 CUDA: fix variable name conflict for Windows build (llama/9382) Johannes Gäßler 2024-09-09 14:22:53 +02:00