Commit Graph

  • 6576c4da90 hexagon: use DIRID 13 in libggml-htp.inf for modern InfVerif (llama/22306) Mengsheng Wu 2026-04-25 00:21:33 +08:00
  • 07d6db39e5 metal : print GPU description (llama/22318) Georgi Gerganov 2026-04-24 13:56:03 +03:00
  • dfb8b68799 ggml : minor coding style (llama/22308) Georgi Gerganov 2026-04-24 11:02:00 +03:00
  • 23921d5a69 hexagon: add SOLVE_TRI op (llama/21974) Mengsheng Wu 2026-04-24 09:39:13 +08:00
  • 641998f558 fix(shader): handle the buffer aliasing for rms fuse (llama/22266) Chen Yuan 2026-04-23 19:32:59 -04:00
  • 71b1ab3784 hexagon: add support for basic and extended Op profiling (llama/22269) Max Krasnyansky 2026-04-23 14:17:21 -07:00
  • 682ee99305 metal : fix event synchronization (llama/22260) Georgi Gerganov 2026-04-23 08:22:49 +03:00
  • 1aba061737 ggml-base: use MATH_LIBRARY variable instead of hardcoded 'm' (llama/22239) Georgi Gerganov 2026-04-23 08:22:08 +03:00
  • b938c5026c sycl : fused MoE mul_mat_vec_q for TG (llama/21920) abotsis 2026-04-22 23:18:56 -06:00
  • df528c4f71 ggml-webgpu: add support for im2col (llama/22259) Chen Yuan 2026-04-22 23:17:41 -04:00
  • b6b547885c CUDA: fuse relu + sqr (llama/22249) Anav Prasad 2026-04-23 02:28:56 +00:00
  • 393fdffe20 HIP: flip GGML_HIP_GRAPHS to default on (llama/22254) uvos 2026-04-23 02:34:31 +02:00
  • d2a26dc8e2 Implement async tensor api and event api (llama/22099) Nikhil Jain 2026-04-22 10:52:01 -07:00
  • 0fbe4c4ca7 ggml-webgpu: Add fused RMS_NORM + MUL (llama/21983) Masashi Yoshimura 2026-04-23 02:51:40 +09:00
  • c5bb7c0078 sycl: Improve mul_mat_id memory efficiency and add BF16 fast path (llama/22119) Akarshan Biswas 2026-04-22 18:02:56 +05:30
  • 447be522e9 ggml-webgpu(shader): support conv2d kernels. (llama/21964) Chen Yuan 2026-04-21 23:18:57 -04:00
  • d6a417408c hexagon: add support for FILL op (llama/22198) Aparna M P 2026-04-22 04:54:20 +05:30
  • 2e5eb6e951 ggml-webgpu: reset CPU/GPU profiling time when freeing context (llama/22050) Masashi Yoshimura 2026-04-22 08:05:21 +09:00
  • 84a6b5c039 Hexagon: DAIG op (llama/22195) Shreya Jain 2026-04-21 14:16:04 -07:00
  • e2014d6959 hexagon: fix missing v79 entry in libggml-htp.inf (llama/22194) Mengsheng Wu 2026-04-22 04:53:44 +08:00
  • 3a73f9cf0b openvino: driver setup, CI split, thread safety, and NPU optimizations (llama/21944) Zijun Yu 2026-04-21 23:58:34 +08:00
  • 150cef5a5f metal : workaround macOS GPU interactivity watchdog (llama/22216) Georgi Gerganov 2026-04-21 17:24:55 +03:00
  • 85bbc82209 vulkan: Support F16 OP_FILL (llama/22177) Jeff Bolz 2026-04-21 11:01:56 +02:00
  • e7cffdbd0b ggml : bump version to 0.10.0 (ggml/1463) Georgi Gerganov 2026-04-21 11:02:56 +03:00
  • b13deaabae ggml-cuda: flush legacy pool on OOM and retry (llama/22155) leonardHONG 2026-04-21 05:30:38 +08:00
  • 239c5c86c3 Tensor-parallel: Fix delayed AllReduce on Gemma-4 MoE (llama/22129) Gaurav Garg 2026-04-20 21:55:39 +05:30
  • 6429023e5f TP: fix 0-sized tensor slices, AllReduce fallback (llama/21808) Johannes Gäßler 2026-04-20 18:09:39 +02:00
  • 2b9fb0be77 ggml-cpu: Optimized x86 and generic cpu q1_0 dot (follow up) (llama/21636) pl752 2026-04-20 21:02:54 +05:00
  • 5f21fdcbb9 ggml-webgpu: updated matrix-vector multiplication (llama/21738) neha-ha 2026-04-20 07:37:17 -07:00
  • 931cf2f3a8 Fix reorder MMVQ assert on unaligned vocab sizes (llama/22035) Katostrofik 2026-04-20 01:39:45 -04:00
  • b8f57c9c50 CUDA: refactor mma data loading for AMD (llama/22051) Johannes Gäßler 2026-04-19 18:26:59 +02:00
  • 945746b40c HIP: Remove unesscary NCCL_CHECK (llama/21914) uvos 2026-04-19 12:59:44 +02:00
  • 671fd1527a ggml : reduce CPU overhead in meta backend (llama/22041) Gaurav Garg 2026-04-19 15:18:35 +05:30
  • 171f037fba cmake: remove CMP0194 policy to restore MSVC builds (llama/21934) texasich 2026-04-19 02:25:05 -05:00
  • 32789b9e07 rpc : refactor the RPC transport (llama/21998) Radoslav Gerganov 2026-04-19 10:21:53 +03:00
  • a899e4bdcb ggml-backend-meta: add multi-segment read support in get_tensor (llama/22063) SamareshSingh 2026-04-18 03:04:51 -05:00
  • cbbe935765 ggml-webgpu: fix compiler warnings and refactor FlashAttention encoding (llama/21052) Reese Levine 2026-04-17 09:17:11 -07:00
  • 918e0ad209 CUDA: use LRU based eviction for cuda graphs (llama/21611) Aman Gupta 2026-04-17 23:24:21 +08:00
  • 77c0630ce6 opencl: refactor q8_0 set_tensor and mul_mat host side dispatch for Adreno (llama/21938) lhez 2026-04-16 22:28:33 -07:00
  • b25d5d050b hexagon: optimize HMX matmul operations (llama/21071) nullname 2026-04-17 04:48:34 +08:00
  • 57a48a4850 opencl: add q5_K gemm and gemv kernels for Adreno (llama/21595) shaofeiqi 2026-04-16 12:08:33 -07:00
  • 820438ae2c ggml: add graph_reused (llama/21764) Aman Gupta 2026-04-16 17:21:28 +08:00
  • 655c0750f5 metal: Implement ROLL op (llama/21946) Kusha Gharahi 2026-04-16 03:54:37 -05:00
  • 94d6d0b743 ggml-cpu: add 128-bit RVV implementation for Quantization Vector Dot (llama/20633) rehan-10xengineer 2026-04-16 13:15:15 +05:00
  • 07c181b57f ggml : implemented simd_gemm kernel for riscv vector extension (llama/20627) rehan-10xengineer 2026-04-16 13:14:26 +05:00
  • 092330b474 ggml-webgpu: compute pass batching and removing profiling overhead (llama/21873) Reese Levine 2026-04-16 01:12:19 -07:00
  • f62bb13320 Fix Q8_0 reorder: garbage on 2nd prompt + crash on full VRAM (llama/21638) Katostrofik 2026-04-16 01:34:05 -04:00
  • 7fe6b8e171 vulkan: optimize im2col (llama/21713) Ruben Ortlam 2026-04-15 19:04:51 +02:00
  • c6d1fbf31f cuda: Q1_0 initial backend (llama/21629) Pasha Khosravi 2026-04-15 09:38:38 -07:00
  • 2a785c5969 ggml-webgpu: Fix dequantization helpers to not pass in pointers (llama/21872) Reese Levine 2026-04-15 09:14:40 -07:00
  • 9638e29657 CUDA: require explicit opt-in for P2P access (llama/21910) Johannes Gäßler 2026-04-15 16:01:46 +02:00
  • 7e57b20d53 CUDA: manage NCCL communicators in context (llama/21891) Johannes Gäßler 2026-04-15 15:58:40 +02:00
  • 182db04cb2 rpc : add native RDMA transport for RPC backend (RoCEv2) (llama/20590) Valeriy Dubov 2026-04-15 16:44:02 +03:00
  • 86d94cd95b docs: more extensive RoPE documentation [no ci] (llama/21953) Xuan-Son Nguyen 2026-04-15 14:45:16 +02:00
  • 24cc89e477 hexagon: optimization for HMX mat_mul (llama/21554) Yiwei Shao 2026-04-14 14:09:03 -07:00
  • 44d86c4921 ggml : remove ggml-ext.h (llama/21869) Xuan-Son Nguyen 2026-04-14 16:32:58 +02:00
  • 08e412c862 metal : fix FA support logic (llama/21898) Georgi Gerganov 2026-04-14 17:32:29 +03:00
  • 45365fa111 vulkan: Programmatically add RoundingModeRTE to all shaders when the device supports it (llama/21572) Jeff Bolz 2026-04-14 15:17:45 +02:00
  • 7024f7e5c1 ci : re-enable mac workflows (llama/21894) Georgi Gerganov 2026-04-14 15:58:09 +03:00
  • 691b1d0826 metal : add XIELU unary op (llama/20802) Seyoung Jeong 2026-04-14 21:43:59 +09:00
  • 80f7be74bb ggml : fix ARM NEON nvfp4 dot product on non-dotprod targets (llama/21559) Richard Davison 2026-04-14 13:23:45 +02:00
  • bfdcd4a92c cmake: fix CMP0194 warning on Windows with MSVC (llama/21630) texasich 2026-04-14 05:47:56 -05:00
  • b732f4d9b5 ggml-webgpu: Update register tiling matmul to use f32 accumulation (llama/21644) Reese Levine 2026-04-14 03:46:41 -07:00
  • cdeaa34174 vulkan: Support GGML_TYPE_NVFP4 (llama/21455) Jeff Bolz 2026-04-14 11:34:23 +02:00
  • 0f99a47177 vulkan: Flash Attention DP4A shader for quantized KV cache (llama/20797) Ruben Ortlam 2026-04-13 14:21:31 +02:00
  • d9ed371c2c CUDA: Limit DeviceSegmentedSort to immediate mode (llama/21718) Oliver Simons 2026-04-13 11:14:06 +02:00
  • 36b7bb3d95 Remove extra conditional check on debug mode. (llama/21798) Masashi Yoshimura 2026-04-13 12:13:04 +09:00
  • 655072cd78 sycl: disable Q1_0 in backend and cleanup unused variables (llama/21807) Akarshan Biswas 2026-04-13 07:14:58 +05:30
  • b907207312 mtmd: add Gemma 4 audio conformer encoder support (llama/21421) Stephen Cox 2026-04-13 00:15:26 +12:00
  • c0b46c2f8f CUDA: skip compilation of superfluous FA kernels (llama/21768) Johannes Gäßler 2026-04-11 18:52:11 +02:00
  • e0c8e505e9 opencl: add basic support for q5_k (llama/21593) shaofeiqi 2026-04-11 01:46:19 -07:00
  • 34381b01c4 ggml : fix a few instances of missing GGML_TYPE_Q1_0 cases (llama/21716) Sigbjørn Skjæret 2026-04-11 08:45:00 +02:00
  • 3af7c879bc CUDA: also store node->src ne/nb for graph equality (llama/21736) Aman Gupta 2026-04-11 10:30:30 +08:00
  • 28ce072f59 hexagon: improved Op queuing, buffer and cache management (llama/21705) Max Krasnyansky 2026-04-10 15:47:43 -07:00
  • 2580cfc703 ggml-webgpu: support non-square subgroup matrix configs for Intel GPUs (llama/21669) Rithik Sharma 2026-04-10 10:52:38 -07:00
  • 3fc738a8c2 ggml-webgpu: address quantization precision and backend lifecycle managment (llama/21521) Chen Yuan 2026-04-10 13:52:01 -04:00
  • 458ad1d93e vulkan: Support Q1_0 (llama/21539) Jeff Bolz 2026-04-10 01:35:27 -05:00
  • 28347201fc CUDA: fuse muls (llama/21665) Aman Gupta 2026-04-10 10:24:09 +08:00
  • c77a33df06 HIP: add CDNA4 (gfx950) architecture support for MI350X/MI355X (llama/21570) andyluo7 2026-04-09 22:13:32 +03:00
  • bb895c843d ggml: backend-agnostic tensor parallelism (experimental) (llama/19378) Johannes Gäßler 2026-04-09 16:42:19 +02:00
  • c4c6e143a7 ggml : check return value of CUB calls used in argsort and top-k (they all return cudaError_t) (llama/21676) fairydreaming 2026-04-09 15:17:11 +02:00
  • f0ee409f7b metal : add missing mm-id specializations for q1_0 (llama/21662) Georgi Gerganov 2026-04-09 10:54:00 +03:00
  • 4598eb080b sycl : add flash-attn support for head size 512 (llama/21654) Akarshan Biswas 2026-04-09 12:06:48 +05:30
  • 1d555510de vulkan: unify type macros to use Vx instead of _VECx (llama/21605) Ruben Ortlam 2026-04-09 07:31:51 +02:00
  • 2c7472939f CUDA: also store node->src->data ptrs for equality check (llama/21635) Aman Gupta 2026-04-09 01:01:56 +08:00
  • 16dd171620 fix: free ctx_copy in ggml_opt_free to plug per-training-session leak (llama/21592) RealOrko 2026-04-08 16:40:15 +01:00
  • e70c0d43f4 webgpu : Query for adapter support when registering WebGPU backend (llama/21579) Reese Levine 2026-04-08 06:08:29 -07:00
  • 15deafa31e metal: Q1_0 backend (llama/21528) Pasha Khosravi 2026-04-08 06:07:47 -07:00
  • fa2eaa433b CUDA: make cuda graphs props check faster (llama/21472) Aman Gupta 2026-04-08 09:05:51 +08:00
  • d91d1e8e6c ggml-cuda: ds_read_b128 for q4_0 and q4_1 mmq kernels (llama/21168) iacopPBK 2026-04-07 21:47:42 +02:00
  • d1456437e1 ggml-webgpu: parameterize submission size and add iOS specific limits (llama/21533) Reese Levine 2026-04-07 10:30:01 -07:00
  • 5ef7aafa06 CUDA: check for buffer overlap before fusing (llama/21566) Aman Gupta 2026-04-08 00:57:04 +08:00
  • f1d2b83db0 ggml : deprecate GGML_OP_ADD1 (llama/21363) Georgi Gerganov 2026-04-07 15:28:27 +03:00
  • 78b4fd85e1 ggml: Vulkan build, Linux -- output error string for errno on fork failure (#20868) (llama/20904) Tom Overlund 2026-04-07 07:54:55 -04:00
  • 18c98ffaf7 vulkan: add FA dequant for q4_1, q5_0, q5_1, iq4_nl (llama/21029) mkoker 2026-04-07 07:41:29 -04:00
  • a1f76fb4cf ggml-cuda : fix CDNA2 compute capability constant for gfx90a (MI210) (llama/21519) Antoine Viallon 2026-04-07 12:18:55 +02:00
  • 1ebf3cafa0 Add Q8_0 reorder optimization (~3x tg speedup on Intel Arc) (llama/21527) PMZFX 2026-04-07 04:12:49 -04:00
  • 9cbc4b3acb ggml-webgpu: Add the support of MUL_MAT_ID (llama/21147) Masashi Yoshimura 2026-04-07 05:08:46 +09:00
  • 0c2fbd4703 ggml: add Q1_0 1-bit quantization support (CPU) (llama/21273) Pasha Khosravi 2026-04-06 11:55:21 -07:00
  • 7b19b94c5d Write an optimized flash_attn_stream_k_fixup kernel (llama/21159) Gaurav Garg 2026-04-07 00:04:29 +05:30