Commit Graph

  • cc0d00f610 vulkan: make mul_mm ALIGNED a spec constant (llama/24689) Jeff Bolz 2026-06-23 07:26:17 -05:00
  • e3c18388c5 vulkan: link ggml-cpu when GGML_VULKAN_CHECK_RESULTS / RUN_TESTS are enabled (llama/24444) Wyatt Caldwell 2026-06-23 03:55:46 -07:00
  • 2955e84ec3 ggml-webgpu: improve MTP inference by using mat-vec path for small batches (llama/24811) Masashi Yoshimura 2026-06-23 17:13:55 +09:00
  • 0badff27cb opencl: q8_0 gemv precision improvement (llama/24923) Shawn Gu 2026-06-22 22:25:21 -07:00
  • 1b3f8434f9 support bf16 on bin_bcast OP and unary OPs (llama/24838) Neo Zhang 2026-06-22 19:09:02 +08:00
  • 556d900cfb fix(hexagon): use padded stride for ssm-conv weights (llama/24470) Guanhuai Zhang 2026-06-21 05:58:49 +08:00
  • 420518d88e ggml : optimize AMX (llama/24806) Adrien Gallouët 2026-06-20 12:43:06 +02:00
  • 75b90fb0ea ggml-webgpu: add adapter toggles for F16 on Vulkan + NVIDIA Masashi Yoshimura 2026-06-20 08:12:32 +09:00
  • adbcfa0d46 mtmd, arg: fix utf8 handling on windows (llama/24779) Xuan-Son Nguyen 2026-06-19 22:28:38 +02:00
  • e3ab311094 examples : fix argument flag for min speech duration in VAD (#3907) QuantiusBenignus 2026-06-26 02:09:03 -04:00
  • 43d78af5be examples : update model names in parakeet-cli README.md [no ci] (#3906) Daniel Bevenius 2026-06-23 09:12:31 +02:00
  • bae6bc02b1 Fix pkgconfig configuration (Nix build failure) (#3894) Nicky Mouha 2026-06-22 23:36:48 -04:00
  • 6edfe6009f include parakeet in build-xcframework.sh (#3899) Naitik Shah 2026-06-22 14:06:51 +04:00
  • 5ed76e9a07 talk-llama : sync llama.cpp Georgi Gerganov 2026-06-19 10:19:58 +03:00
  • 41cf1278c9 sync : ggml Georgi Gerganov 2026-06-19 10:19:47 +03:00
  • f92382f1c2 ggml : bump version to 0.15.2 (ggml/1548) Georgi Gerganov 2026-06-19 10:14:26 +03:00
  • 587f602ea1 ggml-cpu: support K tails in power10 Q8/Q4 MMA matmul (llama/24753) shalinib-ibm 2026-06-19 11:25:38 +05:30
  • 50fd261d0d Ggml/cuda col2im 1d (llama/24417) Pascal 2026-06-18 22:23:01 +02:00
  • 5cbc2153c2 hexagon: support for op-trace (fine-grain tracing of HVX/HMX/DMA events) (llama/24592) Max Krasnyansky 2026-06-18 08:35:02 -07:00
  • 4c4f6eaca7 rename GGML_SYCL_SUPPORT_LEVEL_ZERO (llama/24719) Neo Zhang 2026-06-18 16:18:26 +08:00
  • 69a9798a82 sycl : support MUL_MAT and OUT_PROD with Q1_0 (llama/24721) Neo Zhang 2026-06-18 16:17:37 +08:00
  • 312033789f support OPs: conv_2d, conv_2d_dw, conv2d_transpose (llama/24600) Neo Zhang 2026-06-18 14:40:03 +08:00
  • c39dd2db8e metal : check for BF16 support in concat kernel (llama/24747) Georgi Gerganov 2026-06-18 09:16:06 +03:00
  • 6aff09ac48 ggml-cpu: Conditionally enable power11 backend based on compiler support (llama/24687) shalinib-ibm 2026-06-18 00:15:19 +05:30
  • d911569230 metal : implement rope_back operator (llama/24725) Georgi Gerganov 2026-06-17 20:36:05 +03:00
  • ad19c98240 metal : add f16 and bf16 support for concat operator (llama/24724) Georgi Gerganov 2026-06-17 19:38:55 +03:00
  • 6772827b26 add dev2dev memcpy by SYCL API (llama/24476) Neo Zhang 2026-06-17 22:21:34 +08:00
  • 01b1c3ada7 Add conv_3d (llama/24691) Neo Zhang 2026-06-17 22:20:01 +08:00
  • ae09736e83 vulkan: record actual memory properties during buffer creation (llama/24326) Winston Ma 2026-06-17 17:14:48 +08:00
  • 201f69ce6a Revert "cuda: reset cuda context after reading memory size (llama/23935)" (llama/24715) Ruben Ortlam 2026-06-17 10:59:35 +02:00
  • 055bf5c73c ci: fix vulkan docker images (llama/24595) kononnable 2026-06-17 09:43:45 +02:00
  • 9f785839a3 opencl: optimize mul_mat_f16_f32_l4 for decode (llama/24504) lhez 2026-06-16 23:21:26 -07:00
  • 5fa14e9931 openvino: OV 2026.2, context-shift, Q5_1 support, gemma4 dense/embedding, and -fa off (llama/24503) Zijun Yu 2026-06-17 14:11:21 +08:00
  • da66f048b1 sycl : Enable to support fp16 by OPs: SQR, SQRT, LOG, SIN, COS, CLAMP (llama/24692) Neo Zhang 2026-06-17 13:58:03 +08:00
  • fddcda58a3 SYCL: fix use-after-free bug with async memcpy in MoE prefill (llama/24676) Alexey Kopytko 2026-06-17 14:57:29 +09:00
  • dd1a6ca897 sycl: Add optional USM system allocations (llama/22526) Francois Dugast 2026-06-17 07:54:21 +02:00
  • 694579182f vulkan: prefer host-visible memory buffers on UMA devices (llama/22930) Winston Ma 2026-06-16 15:36:52 +08:00
  • 93c02083bd vulkan: Support gated_delta_net with S_v=16 (llama/24581) Jeff Bolz 2026-06-16 02:26:57 -05:00
  • fa204c82f4 sycl: support reordered Q4_K/Q5_K/Q6_K MoE MUL_MAT_ID (llama/24452) Frosty40 2026-06-16 00:35:00 -05:00
  • 79f88a1104 Support OP EXPM1, support all UT cases of FLOOR, TRUNC, ROUND (llama/24363) Neo Zhang 2026-06-16 13:34:29 +08:00
  • c8f370a460 vulkan: add col2im_1d op (llama/24425) Pascal 2026-06-16 06:34:43 +02:00
  • d77b2f704c vulkan: support more CONCAT types (llama/24579) Jeff Bolz 2026-06-15 06:19:21 -05:00
  • bd3912b0a8 wasm : fix fallback symbol collision (llama/24639) Andrei 2026-06-15 00:11:59 -07:00
  • 1fd857e2f9 SYCL: use native subgroup size for K-quant DMMV (llama/21700) Katostrofik 2026-06-15 03:10:53 -04:00
  • 96051f04a7 sycl: fix soft_max_f32 max reduction (llama/24451) someoneinjd 2026-06-15 15:10:12 +08:00
  • e958dcead1 sycl : fix reorder function; add fp32/fp16 in build script (llama/24578) Neo Zhang 2026-06-15 15:08:34 +08:00
  • d20057908a sycl : enhance set_rows to support q1_0, mxfp4, nvfp4 (llama/24564) Neo Zhang 2026-06-15 15:01:40 +08:00
  • 3cb087c42a add to support pool_1d, move pool_1d/2d code to pool.cpp/hpp (llama/24584) Neo Zhang 2026-06-15 15:01:07 +08:00
  • 5832e734d4 Remove per-allocation Level Zero runtime checks (llama/23399) Alexey Kopytko 2026-06-15 15:58:42 +09:00
  • 1b2d6d2c23 metal : add repeat bf16 (llama/24638) Georgi Gerganov 2026-06-15 09:57:16 +03:00
  • 7349e5ae11 CUDA: only support F32/F16 for GGML_OP_REPEAT (llama/24533) leonardHONG 2026-06-15 14:11:00 +08:00
  • 3e0b917514 ggml-webgpu: improve i-quants mul_mat performance and speed up prefill (llama/24530) Masashi Yoshimura 2026-06-15 10:15:30 +09:00
  • 1216e0957b vulkan: support non-contig unary/glu ops (llama/24215) Jeff Bolz 2026-06-13 08:44:15 -05:00
  • dc195118ef vulkan: add pipeline barriers for memcpy read operations (llama/23770) Ruben Ortlam 2026-06-12 16:43:50 +02:00
  • f049fff95a release : v1.9.1 (#3892) v1.9.1 Daniel Bevenius 2026-06-19 06:12:37 +02:00
  • 200b119790 ci : add GGML_NATIVE=OFF and GGML_BMI2=OFF to windows-blas (#3891) Daniel Bevenius 2026-06-18 14:49:08 +02:00
  • 86c40c3bd6 release : v1.9.0 (#3886) v1.9.0 Daniel Bevenius 2026-06-17 11:36:57 +02:00
  • 0d14756929 ruby : add support for Parakeet (#3885) KITAITI Makoto 2026-06-17 13:42:09 +09:00
  • 9efddafb91 parakeet : add support for NVIDIA Parakeet (#3735) Daniel Bevenius 2026-06-16 20:44:10 +02:00
  • 3805e602d3 ci : only trigger release jobs for tags (#3883) Daniel Bevenius 2026-06-16 14:33:42 +02:00
  • 48f628a848 release : v1.8.7 (#3881) v1.8.7 Daniel Bevenius 2026-06-16 12:28:23 +02:00
  • db5a84bd79 cli : add --version flag (#3878) Rum Nguyen 2026-06-16 13:58:09 +07:00
  • 0ec0845110 talk-llama : sync llama.cpp Georgi Gerganov 2026-06-15 09:15:48 +03:00
  • 0a3fa9ca17 sync : ggml Georgi Gerganov 2026-06-15 09:13:43 +03:00
  • f35f47b5d2 ggml : bump version to 0.15.1 (ggml/1541) Georgi Gerganov 2026-06-12 15:32:00 +03:00
  • 882736f886 ggml: support concat for scalar types at cuda backend (llama/24011) ZihaoMu 2026-06-12 14:32:44 +08:00
  • 2dcfd49d59 opencl: add q5_0/q5_1 gemm and gemv kernels for Adreno (llama/24319) shaofeiqi 2026-06-11 21:43:09 -07:00
  • afd559279c vulkan: ifdef eMesaHoneykrisp (build fix) (llama/24479) Jeff Bolz 2026-06-11 13:22:17 -05:00
  • b04008fcec ggml : bump version to 0.15.0 (ggml/1539) Georgi Gerganov 2026-06-11 19:32:38 +03:00
  • 6870cfd616 vulkan: add fast path for contiguous buffer transfers (llama/23973) Winston Ma 2026-06-11 21:46:25 +08:00
  • a512e4c5c3 vulkan: use medium matmul tile on Asahi Linux (llama/24306) Kevin Liu 2026-06-11 09:43:04 -04:00
  • 1a1900f90c Remove padding and multiple D2D copies for MTP (llama/24086) Gaurav Garg 2026-06-10 23:21:16 +05:30
  • ef85b26d9f CUDA: Fix ssm_scan_f32 data-races (llama/24360) Oliver Simons 2026-06-10 14:27:08 +02:00
  • dc794303d8 vulkan: reduce iq1 shared memory usage for mul_mm (llama/24287) Jeff Bolz 2026-06-09 06:27:38 -05:00
  • 686bc802d1 vulkan: add v_dot2_f32_f16 support in matrix-matrix multiplication and Flash Attention (llama/24123) Ruben Ortlam 2026-06-09 13:27:04 +02:00
  • 28c7ed3db7 ggml : add GGML_OP_COL2IM_1D (llama/24206) Pascal 2026-06-09 11:01:37 +02:00
  • 2d68a3066f ggml-cpu : fix rms_norm_back wrong output under in-place aliasing (llama/24305) Yash Raj Pandey 2026-06-09 03:24:27 -04:00
  • 72894aa250 Remove case for GGML_TYPE_Q4_K in mvvq.cu (llama/23528) ravel7524 2026-06-09 07:46:23 +02:00
  • e69e5138fe ggml-webgpu: Add clang-format job (llama/24308) Reese Levine 2026-06-08 20:54:24 -07:00
  • aa42b48312 ggml-webgpu: Improve prefill speeds for k-quants + refactor matmul for Q4/Q5/Q8 and k-quants (llama/24225) Masashi Yoshimura 2026-06-09 07:19:56 +09:00
  • 15e5d401d1 Handle buffer overlap / buffer aliasing for concat operator (llama/24000) Nikhil Jain 2026-06-08 08:07:31 -07:00
  • 490e50056c Implement 2D workgroups for scale, binary, and unary ops (llama/24044) Nikhil Jain 2026-06-08 08:07:15 -07:00
  • fbf720dc9f vulkan: Use cm2 decode_vector for mul_mat_id B matrix loads (llama/23991) Jeff Bolz 2026-06-08 03:40:37 -05:00
  • 782f1226c8 cuda: reset cuda context after reading memory size (llama/23935) Ruben Ortlam 2026-06-08 10:22:44 +02:00
  • df7638d822 ci : pin github actions to commit sha's (#3865) Daniel Bevenius 2026-06-09 12:51:00 +02:00
  • ba573929cd coreml : fix --quantize crash for mlprogram format; fix --optimize-ane label (#3868) Christopher Albert 2026-06-09 08:34:31 +02:00
  • 84bd03a438 talk-llama : sync llama.cpp Georgi Gerganov 2026-06-08 12:55:06 +03:00
  • 4df9a57df2 sync : ggml Georgi Gerganov 2026-06-08 12:52:27 +03:00
  • b31466b4a1 ggml : bump version to 0.14.0 (ggml/1533) Georgi Gerganov 2026-06-08 12:51:59 +03:00
  • b932ec5529 sync : ggml Georgi Gerganov 2026-06-08 12:52:17 +03:00
  • 4669631d20 HIP: add gfx1152 and gfx1153 to RDNA3.5 (llama/24129) Harkirat Gill 2026-06-08 02:33:23 -04:00
  • 2c139c2e5e metal : fix im2col 1D case (audio models) (llama/24220) Xuan-Son Nguyen 2026-06-08 08:03:18 +02:00
  • 1777deff4c vulkan: check coopmat2 features before reporting support (llama/24186) Ruben Ortlam 2026-06-06 09:11:35 +02:00
  • a87e950a06 opencl: improve get_rows, cpy, concat and q6_k flat gemv (llama/24160) lhez 2026-06-05 13:45:25 -07:00
  • 5a1feed8ca vulkan: add fwht support for Intel with shmem reduction (llama/23964) Ruben Ortlam 2026-06-05 19:44:40 +02:00
  • facb02c4c3 kleidiai : dynamic chunck-based scheduling for hybrid execution (llama/23819) Charles Xu 2026-06-05 09:11:47 +02:00
  • 4fa1e0687e CUDA: enroll mul_mat_vec_q_moe into pdl (llama/24087) Oliver Simons 2026-06-05 08:37:34 +02:00
  • 4ecede8c8b sycl : port multi-column MMVQ from CUDA backend (llama/21845) Mason Milburn 2026-06-05 01:10:31 -04:00
  • 991b5a8b4a ggml: vectorize ggml_vec_dot_q4_1_q8_1 with WASM SIMD128 (llama/22209) Kartik Sirohi 2026-06-04 18:42:38 +05:30
  • 9d6e561f69 metal : reduce rset heartbeat from 500ms -> 5ms (llama/24074) Georgi Gerganov 2026-06-04 08:05:32 +03:00