Commit Graph

  • dfba90db63 webui: parse effective-parameter sizes (E2B, E4B) as params (#25529) Emanuil Rusev 2026-07-14 18:12:22 +03:00
  • 00e79f6fb1 opencl: fix a dp4a bug for devices where cl_khr_integer_dot_product is unavailable (#25639) b10007 Hongqiang Wang 2026-07-14 08:08:13 -07:00
  • 17a05e451f ui: fix mcp panel for toggle + timeout + proxy + ON/OFF state (#25631) Pascal 2026-07-14 16:50:44 +02:00
  • 7f575c39d6 DeepseekV4: fix seq_rm (#25588) b10005 Aman Gupta 2026-07-14 21:45:36 +08:00
  • 234252e03d label virtual devices in description; add GPUx2 server CI jobs ci/cuda-virtual-devices Anav Prasad 2026-07-14 13:45:17 +00:00
  • 7cbd61002d vulkan/cpu: Support f16 as SET_ROWS src. (#25432) b10004 Jeff Bolz 2026-07-14 08:26:55 -05:00
  • 8ff8c4299d tokenize : align usage by using common args (#25516) b10003 Adrien Gallouët 2026-07-14 15:20:53 +02:00
  • a7312ae94f ggml : add a set of functions for checking contiguity of inner tensor dimensions (#25650) b10002 fairydreaming 2026-07-14 14:37:52 +02:00
  • 6554db6ed0 add link to PR xsn/server_cors_args Xuan Son Nguyen 2026-07-14 13:07:03 +02:00
  • 9562fea713 fix test Xuan Son Nguyen 2026-07-14 12:58:13 +02:00
  • e1fa323932 add tests Xuan Son Nguyen 2026-07-14 12:28:32 +02:00
  • 941011dabb add special "localhost" value Xuan Son Nguyen 2026-07-14 12:21:52 +02:00
  • 657e01125a tests: export-graph-ops: exit gracefully when called w/o arguments (#25619) b10001 Christian Kastner 2026-07-14 12:15:41 +02:00
  • 47a39665e7 ggml: uniformize im2col dst_type for all conv ops (#23660) b10000 JusteLeo 2026-07-14 12:13:13 +02:00
  • 47c786924a kleidiai : add SME2 f32 kernel (#24414) b9999 Charles Xu 2026-07-14 12:12:18 +02:00
  • c9330ed0cf ui: add reasoning effort control to mobile add sheet (#25539) Pascal 2026-07-14 12:05:40 +02:00
  • 8dff020c64 server: add --cors-* options Xuan Son Nguyen 2026-07-14 12:05:05 +02:00
  • cb489bc0fb convert_hf_to_gguf: support split MTP export for HY V3 (#25641) Thiago Padilha 2026-07-14 06:43:15 -03:00
  • ec0dbef816 arg: Flush log before exiting after usage() (#25504) b9996 Christian Kastner 2026-07-14 11:03:22 +02:00
  • c1063ac9d7 sycl: set fattn_vec_nthreads to 256 for Battlemage (#25205) b9995 Titaniumtown 2026-07-14 05:00:00 -04:00
  • 14d3ba45f3 metal : add Q2_0 support (#25419) b9994 Pasha Khosravi 2026-07-13 21:52:00 -07:00
  • 2969d6d15d model: add Hy3 (hy_v3) support with MTP speculative decoding (#25395) b9993 Satinder Grewal 2026-07-14 10:31:04 +12:00
  • 6eddde06a4 CUDA: refactor MMQ kernel configuration (#24127) b9992 Johannes Gäßler 2026-07-13 18:37:57 +02:00
  • e920c523e3 vulkan: Use native e2m1 and e4m3 conversions for mxfp4/nvfp4 (#25338) Jeff Bolz 2026-07-13 08:44:17 -05:00
  • 259ae1df8b spec: add Minimax2 eagle3 support b9990 Adrian 2026-07-13 06:22:37 -07:00
  • 4193ea697f readme : add link to maintainer PRs (#25621) Georgi Gerganov 2026-07-13 16:07:58 +03:00
  • f4253ef965 tests: Harmonize header use (#25616) b9988 Christian Kastner 2026-07-13 14:36:51 +02:00
  • ad8d821991 gguf : add tensor shape accessor (#24405) b9987 QuintinShaw 2026-07-13 18:55:15 +08:00
  • 91c631b21d chat : fix reasoning leak with force-opened bare <think> templates (#24674) b9986 Frosty40 2026-07-13 02:45:10 -05:00
  • efb3036c18 sycl: add fused top-k MoE (#25217) b9985 Frosty40 2026-07-13 01:56:41 -05:00
  • e474bba7af sycl: add Q2_K to DMMV reorder path (#25064) b9984 Todd Malsbary 2026-07-12 23:53:39 -07:00
  • 38fd5c9993 ui: Remove recommended MCP Servers + improve MCP Servers Settings UI/UX (#25535) Aleksander Grygier 2026-07-13 08:45:04 +02:00
  • 99f3dc3229 server: honour per-request reasoning_budget_tokens in chat completions (#23116) b9982 Bernard Ladenthin 2026-07-13 01:58:44 +02:00
  • 34558825a2 vendor : update cpp-httplib to 0.50.1 (#25576) b9981 Alessandro de Oliveira Faria (A.K.A.CABELO) 2026-07-12 20:10:03 -03:00
  • 8014d2cf97 server: Don't consider models with --no-mmproj-auto as multimodal (#25590) b9980 Sebastian Dröge 2026-07-13 01:48:13 +03:00
  • 4114ba18b2 mtmd: fix silent prompt truncation on embedded NUL (#25548) b9979 Pascal 2026-07-13 00:47:25 +02:00
  • 0c4fa7a989 server : evict checkpoints within min-step of each other (#25472) b9978 Aldehir Rojas 2026-07-12 15:59:14 -05:00
  • 6b4dc2116a server : fix image blocks in tool_result being dropped during Anthropic OpenAI conversion (#22536) b9977 quei 2026-07-12 23:43:51 +08:00
  • e3546c7948 Fix conditional to display 'LLAMA_SPLIT_MODE_TENSOR not implemented for architecture' message (#24926) b9976 kdkd 2026-07-11 13:03:24 -05:00
  • d72bfa38f7 gguf : reject empty metadata keys (#24917) b9975 Rohit Mahesh 2026-07-11 13:02:44 -05:00
  • 3cec3bcd16 cuda: Don't crash when querying memory on device with no free memory. (#25157) b9974 cphlipot 2026-07-11 10:13:43 -07:00
  • 13f2b28b09 DeepseekV4: clear cache only for seq rather than full (#25521) b9973 Aman Gupta 2026-07-11 23:35:45 +08:00
  • c92e806d1c server: allow stream for exec_shell_command (#25526) b9972 Xuan-Son Nguyen 2026-07-11 12:42:55 +02:00
  • ea1f7bbb5d server: refactor server_stream (#25541) b9971 Xuan-Son Nguyen 2026-07-11 12:41:47 +02:00
  • 00f5442cc4 ggml : add GGML_OP_LIGHTNING_INDEXER that implements DeepSeek V3.2/V4 lightning indexer (#24231) b9970 fairydreaming 2026-07-11 11:39:07 +02:00
  • 76f2798059 Vulkan: route large matmuls to medium tile on Adreno (#24877) b9969 Raman Shinde 2026-07-11 13:58:29 +05:30
  • 1d1d9a9ed7 opencl: add int8 dp4 dense and MoE prefill optimization for Adreno GPUs (#25537) b9968 Hongqiang Wang 2026-07-10 23:05:58 -07:00
  • 4f37f51972 server: accept null sampling params (#25538) b9967 Pascal 2026-07-10 22:07:29 +02:00
  • c749cb0417 llama : make tensor-split regex patterns static (#24710) b9966 eduardopessin 2026-07-10 18:04:12 +01:00
  • 67776eaee5 hexagon: improve ARGSORT performance for small tensors (#25512) b9965 Max Krasnyansky 2026-07-10 09:06:06 -07:00
  • 22b69b6e92 arg: prevent duplicate spec model downloads (#25527) b9964 Xuan-Son Nguyen 2026-07-10 16:53:26 +02:00
  • 3e706dd55f mtmd: deepseek-ocr v1 multi-tile (#24717) b9963 Xuan-Son Nguyen 2026-07-10 16:05:49 +02:00
  • 3203ab47a6 server : prevent LRU eviction of models still loading gg/server-fix-lru-evict-loading Georgi Gerganov 2026-07-10 16:49:57 +03:00
  • 1f66d544cb Merge branch 'master' into xsn/mtmd_ds_ocr_tiles xsn/mtmd_ds_ocr_tiles Xuan Son Nguyen 2026-07-10 15:11:44 +02:00
  • 83c7477b3f test-deepseek-ocr: relax v1 single-view tolerance; drop trailing prompt space; make DRY opt-in and n_predict model-specific (#25486) Saba Fallah 2026-07-10 15:10:58 +02:00
  • 07d9378286 feat: pre-select models in the webui using alias (#25492) felix 2026-07-10 13:04:00 +00:00
  • 9f623c683d ui: use server modalities in non-router mode (#24874) Josh Leverette 2026-07-10 08:03:52 -05:00
  • a935fbffe1 server: remove loading.html (#25500) b9960 Xuan-Son Nguyen 2026-07-10 14:42:17 +02:00
  • 65dd9133ba metal : add CONV_2D_DW (depthwise convolution) support (#21565) dev-metal Georgi Gerganov 2026-07-10 14:51:31 +03:00
  • e3ea5dea41 metal : add set_rows with src0 f16 (#25434) Georgi Gerganov 2026-07-08 16:51:31 +03:00
  • de755554b1 metal: add col2im_1d op (f32/f16/bf16) (#25176) Georgi Gerganov 2026-07-08 16:50:17 +03:00
  • 383521467f metal : per-op source split + parallel compile (#24021) YiChen Lv 2026-06-20 18:36:32 +08:00
  • 0badc06ab5 sync : ggml b9959 Georgi Gerganov 2026-07-10 13:10:49 +03:00
  • ac17f8ac1c ggml : use ggml_vqtbl1q_u8 for 32-bit compat (whisper/0) Georgi Gerganov 2026-07-10 11:06:42 +03:00
  • c4ae9a88f8 server: improve tools, remove apply_diff (#25498) b9957 Xuan-Son Nguyen 2026-07-10 11:52:59 +02:00
  • 1b9691bcd5 cli: fix crash on wrong server base url (#25497) b9956 marcoStocchi 2026-07-10 11:52:20 +02:00
  • c7af942e8f ui: prevent tooltip from flickering open and closed on hover (#25503) Pascal 2026-07-10 11:49:52 +02:00
  • 8f114a9b57 sync : ggml (#25517) Georgi Gerganov 2026-07-10 10:28:39 +03:00
  • d46786f296 ui: export full message tree instead of active path only (#25501) Pascal 2026-07-10 09:10:45 +02:00
  • 2ed3c1abbb llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 (#25370) b9952 fairydreaming 2026-07-10 09:06:58 +02:00
  • 082b326fc7 ggml-et: Initial ET backend (#24179) b9951 Martin Chang 2026-07-10 12:38:34 +08:00
  • 961e4b26a7 llama-batch: add unit test (#25471) b9950 Aman Gupta 2026-07-10 11:04:31 +08:00
  • 049326a000 opencl: cluster-parallel decode FA for Adreno (#25473) b9949 Hongqiang Wang 2026-07-09 11:13:48 -07:00
  • 074944998d ggml : process data in smaller chunks in CUDA ggml_top_k() and ggml_argsort() to reduce temporary buffers memory usage (#24776) b9948 fairydreaming 2026-07-09 20:07:12 +02:00
  • 3de7dd4c8f cli: add --output option (#25484) b9947 Xuan-Son Nguyen 2026-07-09 19:37:39 +02:00
  • fb30ba9a6c hexagon: tiling, tracing and optimizations for unary ops (#25474) b9946 Aparna M P 2026-07-09 22:45:47 +05:30
  • 82fce65d8b server : move chat-template thinking probe inside the init try/catch (#24093) b9945 Jesse LaRose 2026-07-09 12:37:39 -04:00
  • 5c3a586860 ggml : fix conv 2d dw (#25490) Georgi Gerganov 2026-07-09 17:56:32 +03:00
  • c15c5c77a4 meta: add hard emphasis on agents not writing descriptions/comments (#25480) Piotr Wilkin (ilintar) 2026-07-09 15:18:07 +02:00
  • f84a519403 Refactor: Consistently use smart pointers in test-backend-ops (#25440) Oliver Simons 2026-07-09 15:00:17 +02:00
  • e5d8b4685c mtmd: dsocr-tiles fixes (#25481) Saba Fallah 2026-07-09 13:45:21 +02:00
  • 7ab9323d1d cli: add --output option xsn/cli_output Xuan Son Nguyen 2026-07-09 13:44:21 +02:00
  • 683f0c72e5 Only index by compile times + always multiply/add (#25445) b9941 Oliver Simons 2026-07-09 13:23:57 +02:00
  • 259f2e2a53 llama-bench : init params.offline (#25476) b9940 Adrien Gallouët 2026-07-09 11:56:56 +02:00
  • 92b187c97e metal : add CONV_2D_DW (depthwise convolution) support (#21565) b9939 Sou-ly 2026-07-09 18:29:15 +09:00
  • ccb0c34223 ggml-hip: enable -funsafe-math-optimizations (#24668) b9938 RapidMark 2026-07-09 01:02:26 -07:00
  • 2021515a1a cuda: align snake fusion matcher with the other backends (#25460) b9937 Pascal 2026-07-09 10:00:06 +02:00
  • 64c8b7db72 server : respect min-step when splitting prompt batches (#25420) b9936 Aldehir Rojas 2026-07-09 01:23:30 -05:00
  • f2d1c2f398 hexagon: add VISION RoPE support (#25216) b9935 Aparna M P 2026-07-09 10:25:00 +05:30
  • 32e41fa5b4 ggml-webgpu: tune subgroup split (d_split) in flash_attn_vec (#25418) b9934 Masashi Yoshimura 2026-07-09 08:34:19 +09:00
  • 92366df30d opencl: Q6_K GEMM/GEMV fix for ne01 of weights that are not multiples of 128. (#25464) b9933 Hongqiang Wang 2026-07-08 15:52:21 -07:00
  • a646006f09 vulkan: disable FA mask_opt on GCN to improve performance (#24362) b9932 Ruben Ortlam 2026-07-08 19:01:25 +02:00
  • 167d057604 opencl: ragged-tile MoE prefill FP16 GEMM optimization (skip padded expert tiles) (#25433) b9931 Hongqiang Wang 2026-07-08 09:44:55 -07:00
  • 1ee093937f llama-batch: fix allowed decreasing pos in a seq (#25449) b9930 Aman Gupta 2026-07-09 00:24:34 +08:00
  • 0bbc87b163 vulkan: for small AMD GPUs, reduce submission threshold based on CU count (#25240) b9929 Ruben Ortlam 2026-07-08 18:15:18 +02:00
  • 81ff7abe50 hexagon: new vtcm layouts and improved pipelines for MUL_MAT, MUL_MAT_ID and FLASH_ATTN_EXT (#25425) b9928 Max Krasnyansky 2026-07-08 07:38:27 -07:00
  • 2976e33c04 disable NCCL path when virtual devices are used Anav Prasad 2026-07-08 13:27:02 +00:00
  • a8eca04c3f support cuda virtual devices Anav Prasad 2026-06-25 10:26:56 +00:00
  • c264f65ff9 cli : move to HTTP-based implementation (#24948) b9927 Xuan-Son Nguyen 2026-07-08 14:52:43 +02:00
  • 07e012afdc Make hip quality check run on all changes (#25403) Oliver Simons 2026-07-08 14:38:51 +02:00