Commit Graph

  • f6f32a7f51 try to fix window cublas CI failure Daniel Bevenius 2026-05-11 14:07:30 +02:00
  • 54ecc9dba4 talk-llama : sync llama.cpp Georgi Gerganov 2026-05-10 17:34:06 +03:00
  • 4730e76552 sync : ggml Georgi Gerganov 2026-05-10 17:27:59 +03:00
  • cf6e65bc59 ggml : bump version to 0.11.1 (ggml/1484) Georgi Gerganov 2026-05-10 16:57:19 +03:00
  • 7072bdab92 internal AllReduce kernel for CUDA provider (llama/22299) scutler-nv 2026-05-10 02:05:22 -07:00
  • 8c7efe885c SYCL: reduce allocation overhead during flash attention (llama/22732) Alexey Kopytko 2026-05-09 15:30:39 +09:00
  • 25f543175d Add BF16 support to GET_ROWS operation (llama/21391) Devedse 2026-05-09 07:50:24 +02:00
  • 3542894544 sycl: Q5_K reorder MMVQ/dequant + Q8_0 reorder MMVQ path (llama/22152) Intel AI Get-to Market Customer Success and Solutions 2026-05-08 22:48:07 -07:00
  • 63f7883206 sycl: Battlemage AOT build via spir64_gen + MMQ subgroup annotations (llama/22147) Intel AI Get-to Market Customer Success and Solutions 2026-05-08 22:42:40 -07:00
  • 197c62c10b Add flash attention MMA / Tiles to support MiMo-V2.5 (llama/22812) AesSedai 2026-05-08 20:28:29 -07:00
  • 42aea65eda hexagon: add HTP kernel for GGML_OP_GATED_DELTA_NET (llama/22837) Yanzhao Wang 2026-05-08 17:12:04 -07:00
  • 892f786a65 sycl: support non-contiguous input in PAD op (llama/22148) Intel AI Get-to Market Customer Success and Solutions 2026-05-08 17:05:22 -07:00
  • e0573051c6 Feature hexagon l2 norm (llama/22816) Pranav Dhinakar 2026-05-08 13:41:40 -07:00
  • 184f1a1383 cuda: fuse snake activation (mul, sin, sqr, mul, add) (llama/22667) Pascal 2026-05-08 11:44:09 +02:00
  • ea459fba9d CUDA: lower-case PCI bus id, standardize for ggml (llama/22820) Johannes Gäßler 2026-05-08 10:09:38 +02:00
  • 803424ac5a vulkan: fix spv shadowing (llama/22760) miyan 2026-05-08 15:35:22 +08:00
  • eb38a02de1 ggml: update SCHED_DEBUG output to use ggml_op_desc() (llama/22825) Max Krasnyansky 2026-05-07 22:43:04 -07:00
  • ef77e10404 opencl: add q4_0 MoE GEMM for Adreno (llama/22731) Shawn Gu 2026-05-07 21:17:07 -07:00
  • 6e91ed3b33 CUDA: batch out_prod inner loop with cublasSgemmStridedBatched (llama/22651) leonardHONG 2026-05-08 03:59:29 +08:00
  • 5fd75cda3f llama : fix device state save/load (llama/22805) Georgi Gerganov 2026-05-07 21:43:40 +03:00
  • 7774fe2c8d opencl: add opfilter regex for debugging (llama/22782) shaofeiqi 2026-05-07 11:00:20 -07:00
  • bd693bb1eb sycl: add FILL, CUMSUM, DIAG, SOLVE_TRI, SSM_SCAN, GATED_DELTA_NET (llama/22149) Intel AI Get-to Market Customer Success and Solutions 2026-05-07 08:51:33 -07:00
  • 4395364605 ggml-cpu: Optimized risc-v cpu q1_0 dot pl752 2026-05-07 18:09:25 +05:00
  • d3f16afcf5 ggml-cpu: fuse RMS_NORM + MUL on CPU backend (llama/22423) zzzzwc 2026-05-06 15:41:14 +08:00
  • 3613268bc7 ggml : use CL_DEVICE_GLOBAL_MEM_SIZE as memory estimate for OpenCL --fit (llama/22688) fl0rianr 2026-05-06 07:12:48 +02:00
  • a6d678954a Hexagon: Process M-tail rows on HMX instead of HVX (llama/22724) Trivikram Reddy 2026-05-05 11:43:03 -05:00
  • f83b6bdc44 opencl: refactor Adreno q4_0 (llama/22335) lhez 2026-05-10 14:52:20 +03:00
  • 0bafd810b6 rpc : use graph uid instead of graph cache (llama/22701) Radoslav Gerganov 2026-05-05 13:47:13 +03:00
  • 716acdb082 ggml : bump version to 0.11.0 (ggml/1478) Georgi Gerganov 2026-05-05 13:14:32 +03:00
  • 6f6103f6d0 llama : add option to save memory in device buffers (llama/22679) Georgi Gerganov 2026-05-05 06:35:07 +03:00
  • 4794432337 ggml : implement fast walsh-hadamard transform for kv rotation (#21352) (llama/22631) Ismail 2026-05-05 04:05:05 +02:00
  • 254f951db8 kleidiai : update to v1.24.0 and use release archive (llama/22549) Charles Xu 2026-05-04 21:13:31 +02:00
  • 36a83b84bb CUDA: use fastdiv for batch index split in get_rows (llama/22650) leonardHONG 2026-05-04 22:24:05 +08:00
  • 0fffe2cdb8 vulkan: delete dead GGML_VK_MAX_NODES def (llama/22621) Atomic-Germ 2026-05-03 22:49:29 -07:00
  • d1d0dc2348 ggml-webgpu: add layer norm ops (llama/22406) Chen Yuan 2026-05-03 23:52:53 -04:00
  • 3bcac0a0c7 fix: CUDA device PCI bus ID de-dupe OOMing (ignoring other 3 gpus entirely) (llama/22533) lucy 2026-05-02 16:19:25 -04:00
  • 9ab94b8cda ggml-virtgpu: fix circular dependency in headers (llama/22557) JusteLeo 2026-05-02 15:28:50 +02:00
  • ff5704a416 opencl: Adreno optimization for MoE - MxFP4 (llama/22301) Shawn Gu 2026-05-01 23:02:24 -07:00
  • 3e9b7d0fef server : fix no_speech_thold not being read (#3783) Andreas Lubbe 2026-05-13 10:37:28 +02:00
  • a604a9b5b0 server: fix params leak between requests (#3784) Andreas Lubbe 2026-05-13 08:54:56 +02:00
  • f08258abd7 whisper : fix max_tokens skipping remaining audio (#3798) annaeina 2026-05-13 13:32:00 +08:00
  • 338cce1e58 server: Add support for controlling token_timestamps directly (#3785) Andreas Lubbe 2026-05-12 07:36:00 +02:00
  • c33c5618b7 whisper : fix incorrect timestamps, usually near silences (#2279) Bjarke Viksøe 2026-05-10 16:24:12 +02:00
  • c81b2dabbc ruby : transcribe without GVL, accept more MemoryViews, Windows support, fix memory size report, improve document (#3775) KITAITI Makoto 2026-05-07 13:28:18 +09:00
  • 4bf733672b talk-llama : sync llama.cpp Georgi Gerganov 2026-05-02 09:01:24 +03:00
  • 18162bcf61 cmake : add FindNCCL.cmake (ggml/0) Georgi Gerganov 2026-05-02 08:54:20 +03:00
  • 8384aa8086 sync : ggml Georgi Gerganov 2026-05-02 08:53:58 +03:00
  • bbdaa21aa7 ggml : remove obsolete rms_norm.wgsl (ggml/0) Georgi Gerganov 2026-05-02 08:51:39 +03:00
  • a5a8496d31 ggml : remove obsoloete wgsl templates (ggml/0) Georgi Gerganov 2026-05-02 08:49:06 +03:00
  • 28f8534532 ggml : bump version to 0.10.2 (ggml/1474) Georgi Gerganov 2026-05-02 08:45:46 +03:00
  • 4861a3eeb5 hexagon: hmx flash attention (llama/22347) Yiwei Shao 2026-05-01 20:29:13 -07:00
  • f2ce24fa5c hexagon: enable non-contiguous row tensor support for unary ops (llama/22574) Aparna M P 2026-05-01 22:39:23 +05:30
  • 9623c1203b ggml-webgpu: Fix vectorized handling in mul-mat and mul-mat-id (llama/22578) Masashi Yoshimura 2026-05-01 23:55:01 +09:00
  • 95053f68e4 vulkan: Support asymmetric FA in coopmat2 path (llama/21753) Jeff Bolz 2026-05-01 15:28:32 +02:00
  • 35cb684129 ggml : try fix win32 build (#0) Georgi Gerganov 2026-05-01 18:53:30 +03:00
  • e10025351c sync : ggml Georgi Gerganov 2026-05-01 13:08:32 +03:00
  • ccd04522f9 ggml-webgpu: add the upscale shader (llama/22419) Chen Yuan 2026-05-01 01:22:18 -04:00
  • b34a9f3d83 ggml-webgpu: Improve performance of mat-vec and mat-mat for MUL_MAT_ID (llama/22464) Masashi Yoshimura 2026-05-01 06:19:10 +09:00
  • 0c7c3ba570 vulkan: add get/set tensor 2d functions (llama/22514) Ruben Ortlam 2026-04-30 17:37:13 +02:00
  • 582d2562a4 CUDA: fix tile FA kernel on Pascal (llama/22541) Johannes Gäßler 2026-04-30 13:04:50 +02:00
  • d74c56862b add fast matmul iquants (llama/22504) Rithik Sharma 2026-04-29 22:58:32 -07:00
  • 66392cf1a2 hexagon: make vmem and buffer-size configurable (llama/22487) Max Krasnyansky 2026-04-29 11:51:21 -07:00
  • aec8e69c2f CUDA: fuse SSM_CONV + ADD(bias) + SILU (llama/22478) Anav Prasad 2026-04-29 11:39:56 -07:00
  • 9f2cec1840 ggml-cpu : disable tiled matmul on AIX to fix page boundary segfault (llama/22293) shalinib-ibm 2026-04-29 16:02:40 +05:30
  • c59a773605 examples : update to Q1_0 Georgi Gerganov 2026-05-01 11:53:27 +03:00
  • 320c048724 sync : ggml Georgi Gerganov 2026-04-30 21:44:28 +03:00
  • ad670182d9 ggml : bump version to 0.10.1 (ggml/1469) Georgi Gerganov 2026-04-29 16:41:45 +03:00
  • 44e7803661 ggml-cuda: refactor fusion code (llama/22468) Aman Gupta 2026-04-29 16:19:33 +08:00
  • 6119537e9a ggml-cpu: cmake: append xsmtvdotii march for SpacemiT IME (llama/22317) qiurui144 2026-04-29 15:59:21 +08:00
  • fa20229eeb ggml-webgpu: Fix bug in FlashAttention support check (llama/22492) Reese Levine 2026-04-29 00:59:00 -07:00
  • 3076725eb0 ggml : add sve tuned code for gemm_q8_0_4x8_q8_0() kernel (llama/21916) hrushitfujitsu 2026-04-29 13:27:37 +05:30
  • 5301139374 TP: fix delayed AllReduce + zero-sized slices (llama/22489) Johannes Gäßler 2026-04-29 08:55:07 +02:00
  • c200b588f8 ggml-cuda: Repost of 21896: Blackwell native NVFP4 support (llama/22196) Michael Wand 2026-04-28 15:47:42 -07:00
  • b553e17071 ggml-cuda: add flash-attn support for DKQ=320/DV=256 with ncols2=32 (… (#22286) lnigam 2026-04-29 01:07:35 +05:30
  • e69c109aac vulkan: Coalesce Q4_K/Q5_K scale loads (llama/21751) Matt Corallo 2026-04-28 15:31:04 +00:00
  • 4ea5b6febc ggml-webgpu: fix buffer aliasing for ssm_scan and refactor aliasing logic (llama/22456) Reese Levine 2026-04-28 07:27:17 -07:00
  • 35fa508360 vulkan: add barrier after writetimestamp (llama/21865) Jeff Bolz 2026-04-28 12:28:12 +02:00
  • 0fa31f9bb6 ggml: improve SPIR-V headers detection with __has_include (llama/21918) Emil Askerov 2026-04-28 13:19:06 +03:00
  • 6fceff2eb4 ggml : skip already registered backends and devices (llama/22296) Adrien Gallouët 2026-04-28 09:02:32 +02:00
  • ca624d86ab ggml : revert to -lm linking instead of find_library (llama/22355) Adrien Gallouët 2026-04-28 08:56:02 +02:00
  • 70e4c0aec0 CANN: add new ops, optimize existing ops (llama/21204) hipudding 2026-04-28 14:27:22 +08:00
  • 9c233f11f0 ggml-webgpu: add Q1_0 support (llama/22374) Rithik Sharma 2026-04-27 15:50:59 -07:00
  • f675a8c926 add fast mat-vec kernels for i-quants (llama/22344) Rithik Sharma 2026-04-27 08:25:45 -07:00
  • c9ba41397c fix: rpc-server cache may not work in Windows environments (llama/22394) unraido 2026-04-27 23:25:09 +09:00
  • f5c3ce17d5 ggml : use 64 bytes aligned tile buffers (llama/21058) Adrien Gallouët 2026-04-27 08:30:55 +02:00
  • 1478450e61 add performance-portable tuning for register-tile and subgroup matmul (llama/22241) Rithik Sharma 2026-04-26 09:26:28 -07:00
  • 7296b9c7fa Fix recurrent state serialization for partial reads and writes (llama/22362) Gaurav Garg 2026-04-26 17:04:40 +05:30
  • 9bf6c3c860 CUDA: better coalesce data-access for contiguous concat (llama/22330) Oliver Simons 2026-04-26 09:21:45 +02:00
  • 2f3df42cdd ggml-cpu : re-enable fast gelu_quick_f16 (llama/22339) Sigbjørn Skjæret 2026-04-26 08:28:14 +02:00
  • 4e11277a19 ggml-cpu: optimize avx2 q6_k (llama/22345) Eve 2026-04-26 06:27:50 +00:00
  • 93a3f37642 opencl: add iq4_nl support (llama/22272) lhez 2026-04-25 21:21:58 -07:00
  • 1be2adf7b3 hexagon: guard HMX clock request for v75+ platforms (llama/22377) Trivikram Reddy 2026-04-25 19:58:26 -05:00
  • da738a74f5 CUDA: reduce MMQ stream-k overhead (llama/22298) Johannes Gäßler 2026-04-25 14:15:03 +02:00
  • 21da84303e metal : optimize Metal Tensor API usage for GGML_OP_MUL_MAT (llama/20962) Developer-Ecosystem-Engineering 2026-04-25 05:14:28 -07:00
  • 6296fd5a90 Optimize Q4_0 mul_mat for Arc770, add scripts (llama/22291) Neo Zhang 2026-04-25 14:20:14 +08:00
  • c235b05d8a ggml-webgpu: support for SSM_SCAN and disable set_rows error checking (llama/22327) Reese Levine 2026-04-24 23:18:15 -07:00
  • c546b0b1bc Hexagon: Bump HMX Frequency to Max Corner (llama/22334) Trivikram Reddy 2026-04-24 15:55:17 -05:00
  • 35d679a4f8 ggml-webgpu: enable FLASH_ATTN_EXT on browser without subgroup matrix (llama/22199) Zheyuan Chen 2026-04-24 10:39:09 -07:00
  • 6576c4da90 hexagon: use DIRID 13 in libggml-htp.inf for modern InfVerif (llama/22306) Mengsheng Wu 2026-04-25 00:21:33 +08:00
  • 07d6db39e5 metal : print GPU description (llama/22318) Georgi Gerganov 2026-04-24 13:56:03 +03:00