Commit Graph

  • c0bc8591e8 hexagon: fix Windows crash when op_poll is enabled (#26029) master adgup-qti 2026-07-23 21:38:10 +05:30
  • f53745b438 add design review xsn/skill_add_models Xuan Son Nguyen 2026-07-23 16:20:08 +02:00
  • 184eb2b901 mention about skills in agents.md Xuan Son Nguyen 2026-07-23 16:00:26 +02:00
  • 2fad89caaf add security review section Xuan Son Nguyen 2026-07-23 15:55:42 +02:00
  • 090cecde1d add codeowners Xuan Son Nguyen 2026-07-23 15:23:30 +02:00
  • ce82869361 nits Xuan Son Nguyen 2026-07-23 15:22:43 +02:00
  • 7afa9586ed add code review skill Xuan Son Nguyen 2026-07-23 15:21:51 +02:00
  • 1e760f9946 add common pitfalls Xuan Son Nguyen 2026-07-23 15:10:14 +02:00
  • 6e7be395c1 skill: add new model Xuan Son Nguyen 2026-07-23 14:58:32 +02:00
  • 1425386fd9 CUDA: fix external compilation of q1_0 MMQ (#25778) Johannes Gäßler 2026-07-23 14:45:51 +02:00
  • e6dd0e29a6 args: refactor mlock/mmap/directio into load-mode (#20834) Aaron Teo 2026-07-23 20:32:56 +08:00
  • 16a2011d40 ggml: fix backend split scheduler race condition 0cc4m/backend-split-sync-fix Ruben Ortlam 2026-07-23 14:07:43 +02:00
  • ef6de50a47 ggml: fix backend split scheduler race condition 0cc4m/vulkan-sync-fixes Ruben Ortlam 2026-07-23 14:07:43 +02:00
  • 76f9383dc6 vulkan: synchronization fixes Ruben Ortlam 2026-07-22 14:05:53 +02:00
  • da296d6e72 contrib: fix leftovers from the AI usage policy update (#26030) Pascal 2026-07-23 12:32:23 +02:00
  • c588c4f476 metal : add f16 type support to leaky relu (#25981) b10103 Ilia Ilmer 2026-07-22 23:45:46 -04:00
  • d941f6e1c9 conversion: fix non-MoE NomicBert GGUF conversion error (#25996) Shahir BIn Zulfiker 2026-07-23 09:01:35 +06:00
  • 4310aa4f87 contrib: allow all AI-generated code in general (#26012) Xuan-Son Nguyen 2026-07-23 00:29:03 +02:00
  • cf512566dc ui: Add a "Default" option for the reasoning selector (#25846) Pascal 2026-07-22 23:09:49 +02:00
  • 4a94f386b9 contrib: allow all AI-generated code in general xsn/contrib_ai_gen Xuan Son Nguyen 2026-07-22 19:35:20 +02:00
  • 1a064ab092 CUDA: Improve NVFP4 W4A4 activation quantization (#25730) Oliver Simons 2026-07-22 19:28:02 +02:00
  • 0278d8362d hexagon: activation ops update (#25974) Todor Boinovski 2026-07-22 09:25:04 -07:00
  • e0833bf686 mtmd: use RAII for setting and resetting non-causal attention (#25723) Niklas Wenzel 2026-07-22 18:10:03 +02:00
  • 61328e6a91 feat(ui): add symbolic math support to JS sandbox via nerdamer (#25948) rankaiyx 2026-07-22 23:52:55 +08:00
  • e8e6c7af24 minor: fix reasoning preserve var for DS4 [no ci] (#25999) Piotr Wilkin (ilintar) 2026-07-22 14:32:54 +02:00
  • 6d5a910c50 common: infer the speculative type from the draft repo sidecars (#25989) Pascal 2026-07-22 13:06:35 +02:00
  • f534da26e4 Fix DeepSeek4 crafted template (#25414) Piotr Wilkin (ilintar) 2026-07-22 12:54:40 +02:00
  • 3ce7da2c85 ggml: enable PowerPC backend variants on AIX (#25983) b10092 shalinib-ibm 2026-07-22 14:56:40 +05:30
  • b4d6c7d8ff ci : fix SYCL package shared library lookup (#25987) b10091 KyleHagy 2026-07-22 02:20:40 -07:00
  • 7347430f44 webgpu : add CONV_2D_DW (depthwise conv2d) kernel (#25847) b10090 m1el 2026-07-22 03:24:44 -05:00
  • c5a4a0bb83 cuda: GET_ROWS quants (#25962) b10089 Pascal 2026-07-22 08:42:47 +02:00
  • 67b9b0e7f6 llama-arch: fix DeepSeek4 APE tensor op (#25945) b10088 helanfxz 2026-07-22 10:55:44 +08:00
  • 1f66c3ce1c Add support for Laguna XS.2 & M.1 (#25165) b10087 Joe Rowell 2026-07-22 03:54:08 +02:00
  • 66e4bf7e59 convert: fix handle HunyuanVL XD-RoPE config (#25514) wendadawen 2026-07-22 06:42:35 +08:00
  • b4aa7dd477 mtmd : use align_corners for qwen3vl vision position embedding interpolation (#25781) b10085 Gerben van V 2026-07-21 23:58:34 +02:00
  • 71102a73f2 hexagon: check tensor type when reusing descriptors (#25968) b10084 Wei Wang 2026-07-22 05:44:22 +08:00
  • 846e991ec3 cuda: add sqrt_softplus in topk-moe for dsv4 (#25896) b10083 Aman Gupta 2026-07-22 00:30:01 +08:00
  • fb0e6b6219 kleidiai : warn once when a weight type has no KleidiAI kernel (#25701) b10082 Kamalesh VS 2026-07-21 21:40:29 +05:30
  • 60f6a17704 common: resolve draft repo to its requested sidecar (#25955) b10081 Pascal 2026-07-21 18:03:43 +02:00
  • fd41bf65a2 server: return 400 instead of 500 on validation error with X-Conversation-Id (#25760) b10080 Pascal 2026-07-21 17:47:54 +02:00
  • 40b740ad05 server : properly handle null llama_context (#25868) b10079 fairydreaming 2026-07-21 17:47:17 +02:00
  • f048010180 vulkan: Refactor vk_queue to use per-instance mutexes and unique handles (#23570) b10078 Winston Ma 2026-07-21 23:40:45 +08:00
  • 5735e10c49 ggml-openvino: Add GGML_BACKEND_DL_IMPL invocation for OpenVINO backend (#25795) b10077 Markus Ebner 2026-07-21 16:43:11 +02:00
  • 305ba519ab CUDA: vectorize same-type get_rows with int4 copy (#25929) b10076 Piotr Wilkin (ilintar) 2026-07-21 15:53:57 +02:00
  • 76f46ad29d hexagon: add CLAMP op (#25934) b10075 Todor Boinovski 2026-07-20 16:12:09 -07:00
  • 2beefef688 ui: Sidebar Conversations Bulk Action + Improved Settings logic/UI (#25815) Aleksander Grygier 2026-07-20 23:40:08 +02:00
  • 91d2fc3875 llama_dsv4: write only used rows in state (#25325) Aman Gupta 2026-07-20 22:43:39 +08:00
  • 4ee6a9af71 ui: fix collapsed user bubble with markdown rendering (#25869) Pascal 2026-07-20 16:28:43 +02:00
  • 43b5e63589 UI: fix Settings/Display tool call content toggle (#25783) Pascal 2026-07-20 16:28:24 +02:00
  • 1521a9ac31 ui: enable the agentic flow when only the JS sandbox is active (#25865) Pascal 2026-07-20 16:22:16 +02:00
  • 178a6c4493 opencl: Support broadcast for Adreno MUL_MAT and honor view_offs for Adreno Q8_0 MUL_MAT for llama-server multi-stream (#25910) b10069 Hongqiang Wang 2026-07-19 22:48:57 -07:00
  • 571d0d540d model: rotate injected K/V cache for DFlash (#25823) b10068 Ruixiang Wang 2026-07-18 15:02:18 +02:00
  • 4937ca83f4 llama-quant : exclude i32 ffn_gate_tid2eid routing table from quantization (#25787) b10067 Yash Raj Pandey 2026-07-18 07:43:18 -04:00
  • 86a9c79f86 opencl: load and use kernel_gemm_moe_q6_k_f32_ns from bin kernel lib (#25797) b10066 lhez 2026-07-17 15:29:29 -07:00
  • 6bdd77f13c opencl: read/write MoE dp4a activation tiles to local memory as 128-bit (vectorized LD/ST perf opt) for Adreno GPUs (#25810) Hongqiang Wang 2026-07-17 12:02:27 -07:00
  • 9f364c765a llama : dsv4 graph fixes gg/ggml-op-offload-fix Georgi Gerganov 2026-07-17 18:43:19 +03:00
  • 9291c45443 ggml : adjust logic for offloading ops to weight's backend Georgi Gerganov 2026-07-17 18:43:08 +03:00
  • 86d86ed439 opencl: transpose q4_K noshuffle scales for coalesced reads (#25805) b10064 Hongqiang Wang 2026-07-17 07:49:43 -07:00
  • 7d56da7e54 sync : ggml b10063 Georgi Gerganov 2026-07-17 16:45:30 +03:00
  • 3727404068 ggml : bump version to 0.17.0 (ggml/1568) Georgi Gerganov 2026-07-17 16:44:55 +03:00
  • 5d5306bf3e tests : initialize all tensors in test_dsv4_hc to avoid NaNs in sentinel tensors (#25822) b10061 fairydreaming 2026-07-17 15:33:35 +02:00
  • bdccb7da38 use minimal shmem size 8 instead of 1 to workaround cm2 compiler bug 0cc4m/vulkan-mul-mm-refactor Ruben Ortlam 2026-07-17 13:21:28 +02:00
  • 2c3ed6086f server : add trace logs for slot prefix similarity computation gg/server-slot-similarity-trace Georgi Gerganov 2026-07-17 13:01:02 +03:00
  • 635cdd5fcc common : auto-download dflash- and eagle3- HF sidecars (#25811) Georgi Gerganov 2026-07-17 12:15:30 +03:00
  • 11fd0a6fb7 ggml-blas: default hadamard mul_mat to cpu routine (#25710) b10059 Aaron Teo 2026-07-17 16:39:33 +08:00
  • b3edd252cf fix unused warning when integer dot glslc support is missing Ruben Ortlam 2026-07-17 09:31:22 +02:00
  • 496b099152 fix missing Q2_0 type Ruben Ortlam 2026-07-17 09:18:18 +02:00
  • 3989780be6 fix compiler warning Ruben Ortlam 2026-07-17 08:08:07 +02:00
  • 0e8431faf9 consolidate shmem tables and reduce size by type spec constant Ruben Ortlam 2026-07-17 07:54:05 +02:00
  • 537ce13df7 fix cm2 bindings Ruben Ortlam 2026-07-16 08:10:30 +02:00
  • 4c58939adb fix cm2 spec constants Ruben Ortlam 2026-07-16 07:33:11 +02:00
  • 000651caaa fix cm2 and shmem init Ruben Ortlam 2026-07-15 15:52:47 +02:00
  • 5954d7108f fix indentation Ruben Ortlam 2026-07-15 14:07:35 +02:00
  • 30a97c211c cleanup Ruben Ortlam 2026-07-15 13:35:32 +02:00
  • a040fbd794 vulkan: use map for mul_mm shapes Ruben Ortlam 2026-07-13 10:56:32 +02:00
  • cde7061b8b vulkan: use spec constant for mul mat type_a Ruben Ortlam 2026-07-07 14:48:02 +02:00
  • 788e07dc91 vulkan: Support Q2_0 (#25430) b10058 Jeff Bolz 2026-07-17 07:42:59 +01:00
  • 0bd0ec6099 sycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (#25690) b10057 Todd Malsbary 2026-07-16 22:49:49 -07:00
  • b85833e934 opencl: add ABS op (#25115) b10056 Gezahegne 2026-07-17 01:13:47 -04:00
  • e8f19cc0ad opencl: loads quants as uint for q4_K and q5_K flat mv (optimization for Adreno A7x GPUs) (#25780) Hongqiang Wang 2026-07-16 13:18:21 -07:00
  • ac2557cb24 docs: added a note about using OpenCl with Adreno 810 (#25786) b10054 akleine 2026-07-16 21:44:45 +02:00
  • 0dc74e332e DeepseekV4: Add fused hyper-connection ops (#25585) Aman Gupta 2026-07-17 00:33:33 +08:00
  • b2dd28a3b6 hexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more MUL_MAT updates (#25762) b10052 Max Krasnyansky 2026-07-16 09:28:04 -07:00
  • f15bd60901 kleidiai: Add SME vs SME2 distinction in kernel dispatch (#25478) b10051 Rajendra Matcha 2026-07-16 21:27:04 +05:30
  • b15ca938ad vulkan: when using transfer queue for async copies, sync on event_wait to avoid race (#25229) b10050 Ruben Ortlam 2026-07-16 15:34:24 +02:00
  • 3278e921b1 conversion: accept BitNetForCausalLM architecture name (#25769) Khashayar Ghafouri 2026-07-16 18:54:47 +05:30
  • 2e1fd76490 TP: fix Phi3, Bert, Plamo2/3, ChatGLM (#25536) b10048 Johannes Gäßler 2026-07-16 15:23:23 +02:00
  • 86b719bf21 vendor: update BoringSSL to 0.20260713.0 (#25624) b10047 Alessandro de Oliveira Faria (A.K.A.CABELO) 2026-07-16 10:17:38 -03:00
  • 32e789fdfd tests: actually exercise test-recurrent-state-rollback (#25758) b10046 Aman Gupta 2026-07-16 21:06:12 +08:00
  • 6eac19fe3e Add tensor name to JSON output cross-profiler Piotr Wilkin 2026-06-06 22:33:01 +02:00
  • 30f64039ca tentative Metal support Piotr Wilkin 2026-05-19 11:52:22 +02:00
  • cdb48f0f14 Add missing unrolls Piotr Wilkin 2026-05-16 15:47:06 +02:00
  • 22bc539f1d Revert accidental change. Piotr Wilkin 2026-05-13 17:22:19 +02:00
  • ca3903d44a Fix braces Piotr Wilkin 2026-05-13 11:09:53 +02:00
  • 85739cfc6f Fix FATTN profiling Piotr Wilkin 2026-05-12 23:58:28 +02:00
  • 4cdd871928 Converge implementation with export-graph-ops Piotr Wilkin 2026-04-07 22:01:00 +02:00
  • 7ba909454b Add missing op parameters to the profiler; add support for test-backend-ops to run performance tests with exactly the tensor shapes from the run Piotr Wilkin 2026-04-03 17:41:57 +02:00
  • d6ab8b3c04 docs, pass copy details Piotr Wilkin 2026-03-29 23:35:38 +02:00
  • 92fd5c2f45 fix mul_mat_id stats, add throughput stat, add envvar trigger, add concurrent mode fix Piotr Wilkin 2026-03-29 22:52:33 +02:00
  • 011b173bc6 fix builds, integrate vulkan profiler, fix copy events, fix export Piotr Wilkin 2026-03-29 16:52:50 +02:00