Commit Graph

  • ed8c26150e cuda : add support for f16->f16 GGML_OP_SET_ROWS (#25367) b9925 fairydreaming 2026-07-08 13:24:20 +02:00
  • 90e0f5cfcb llama: refactor fused ops (#24646) b9924 Aman Gupta 2026-07-08 18:18:09 +08:00
  • bbebeec4a8 server-stream: follow-up on SSE Replay Buffer (#23226) (#25047) b9923 Pascal 2026-07-08 11:02:50 +02:00
  • 230ea9d214 llama-batch: add n_keep_tail in split_equal for recurrent models (#25278) b9922 Aman Gupta 2026-07-08 15:55:19 +08:00
  • f296fdfbed common: auto-create prompts-log-dir at argument parsing, so all tools using the flag benefit (#25322) rankaiyx 2026-07-08 15:45:28 +08:00
  • f1161b15f2 ui: Context usage gauge and panel (#25340) Aleksander Grygier 2026-07-08 09:22:35 +02:00
  • da46e59cbf llama-eval : fix crash when answer is None in HTML dump (#25435) Georgi Gerganov 2026-07-08 10:00:03 +03:00
  • 0512ef1e5a metal : add set_rows with src0 f16 (#25434) b9918 fairydreaming 2026-07-08 08:49:07 +02:00
  • 4a7ee3126d fix: OOB reads in UGM tokenizer (precompiled_charsmap handling) (#18750) b9917 hourhl 2026-07-08 13:02:09 +08:00
  • 57b50e1f6b ggml : fix A indexing in simd_gemm scalar tail-column path (#25390) b9916 tyronecai 2026-07-08 13:00:05 +08:00
  • 68a521b591 ggml : add support for CPU f16->f16 GGML_OP_SET_ROWS (#25344) b9915 fairydreaming 2026-07-08 05:46:28 +02:00
  • 931ca30bef opencl: fix potential crash in aos reconstruct (#25383) b9914 lhez 2026-07-07 20:34:29 -07:00
  • 99b9a6c08a also show model aliases xsn/cli_http_based Xuan Son Nguyen 2026-07-08 00:35:05 +02:00
  • 1729662ea3 nits fixes Xuan Son Nguyen 2026-07-07 22:25:21 +02:00
  • 50ed8076fb no more json in header Xuan Son Nguyen 2026-07-07 22:10:31 +02:00
  • a87b2d77cf pimpl Xuan Son Nguyen 2026-07-07 21:56:44 +02:00
  • b9617e860a cli-view --> cli-ui Xuan Son Nguyen 2026-07-07 21:39:11 +02:00
  • 28b71c022a add ftype Xuan Son Nguyen 2026-07-07 21:36:55 +02:00
  • 7cd7832297 Merge branch 'master' into xsn/cli_http_based Xuan Son Nguyen 2026-07-07 21:11:05 +02:00
  • bec4772f6a Add Q2_0 quantization: type definition and CPU backend (#24448) b9913 Pasha Khosravi 2026-07-07 12:05:47 -07:00
  • c198af4dc2 spec : fix naming, spacing (#25410) b9912 Georgi Gerganov 2026-07-07 18:52:30 +03:00
  • 14f254515f spec : fix naming, spacing gg/spec-naming Georgi Gerganov 2026-07-07 18:43:43 +03:00
  • 3899b39ce2 CUDA: Fuse MMVQ post-scale for NVFP4 (#24481) b9911 Oliver Simons 2026-07-07 17:12:19 +02:00
  • f5525f7e7a server : fix draft model fit vs load inconsistency (#25056) b9910 Alex 2026-07-07 10:20:42 -04:00
  • 5eca4e3cab server : add timings and progress to /responses API stream (#25348) b9909 Thomas LECONTE 2026-07-07 16:13:03 +02:00
  • 6c487e2f79 server: enforce prompt cache RAM limit (#25070) b9908 Thiago Padilha 2026-07-07 10:24:35 -03:00
  • c1a411fb1b common : add missing <fstream> include in common.h (#25220) b9907 zhangrunda 2026-07-07 21:23:53 +08:00
  • 33ca0dcb9d ggml-hip : add -fno-finite-math-only alongside -ffast-math (#25373) b9906 asf0 2026-07-07 05:27:50 -06:00
  • d4a193e269 metal : add set_rows with src0 f16 gg/metal-set-rows-f16 Georgi Gerganov 2026-07-07 14:18:12 +03:00
  • 024c46ae4e llama: fix quantized kv-cache for dsv4 (#25202) b9905 Aman Gupta 2026-07-07 17:46:57 +08:00
  • 108f186d17 [SYCL] fix unsupported UT cases of CONT & CPY (#25231) b9904 Neo Zhang 2026-07-07 17:20:52 +08:00
  • 47e1de77aa [SYCL] support op col2im_1d (#25264) Neo Zhang 2026-07-07 16:07:46 +08:00
  • 55edb2de44 [SYCL] support OP cross_entropy_loss, cross_entropy_loss_back (#25236) b9902 Neo Zhang 2026-07-07 15:48:50 +08:00
  • d209086157 sycl : set K_QUANTS_PER_ITERATION to 1 on DMMV path (#25063) b9901 Todd Malsbary 2026-07-07 00:43:41 -07:00
  • 95e5254c0a [SYCL] fix unsupport ACC UT cases for noncontiguous (#25124) Neo Zhang 2026-07-07 15:40:38 +08:00
  • 9e5ef0dbb1 sycl : enhance argsort to support all UT cases (#25125) b9899 Neo Zhang 2026-07-07 15:39:29 +08:00
  • 3d4cbdf18a sycl : use sycl func to fix AOT double type issue (#25081) b9898 Neo Zhang 2026-07-07 15:38:33 +08:00
  • 26145b3db7 sycl : rename the env vars from "disable" to "enable" (#25042) b9897 Neo Zhang 2026-07-07 15:33:51 +08:00
  • 1a7c25bfdb ggml : make ggml_time_init idempotent (#24422) An Long 2026-07-07 16:29:17 +09:00
  • defa95c306 speculative : fix out-of-bounds read in ngram-map on prompt shrink (#23936) b9895 o7si 2026-07-07 15:25:04 +08:00
  • a8cfdbb9e4 vulkan : check src0 type in GGML_OP_SET_ROWS to avoid failures due to unimplemented f16 support (#25351) b9894 fairydreaming 2026-07-07 06:56:02 +02:00
  • 6f8895feec opencl: general flash attention decode performance optimizations (#25366) b9893 Hongqiang Wang 2026-07-06 19:57:52 -07:00
  • ee445f93d8 common: Set optimal default thread count for ppc ( linux as well as AIX) (#25237) b9892 shalinib-ibm 2026-07-07 03:05:20 +05:30
  • f36e5c348b metal: add col2im_1d op (f32/f16/bf16) (#25176) b9891 Pascal 2026-07-06 20:47:36 +02:00
  • 74976e1aef CUDA: remove -sm row, refactor cuBLAS (#24216) b9890 Johannes Gäßler 2026-07-06 20:04:53 +02:00
  • 9abce7473a server: fix deadlock in load_models() when erasing a finished download (#25358) Pascal 2026-07-06 19:26:06 +02:00
  • cb295bf596 CUDA: extend K-type validation to V-types for flash attention (#24403) b9888 Alexey Kopytko 2026-07-06 23:26:50 +09:00
  • bfdf581b8b server: temporary skip model downloading API test (#25355) Xuan-Son Nguyen 2026-07-06 16:10:04 +02:00
  • 20a04b2206 ggml-cpu: use UE4M3 LUT in ARM NVFP4 dot product (#25331) b9886 ragz4125 2026-07-06 16:36:40 +05:30
  • 3b4fca11ac ggml-cpu: Enable tiled matmul on AIX (#25199) b9885 shalinib-ibm 2026-07-06 15:48:17 +05:30
  • fb5b6b300f chore : replace assert() with GGML_ASSERT() Stanisław Szymczyk 2026-07-06 11:59:38 +02:00
  • d34943ffce ggml : merge ggml_compute_forward_set_rows_f32() and ggml_compute_forward_set_rows_f16() into ggml_compute_forward_set_rows_impl() Stanisław Szymczyk 2026-07-06 11:56:44 +02:00
  • 3965ed4dee ggml : add missing type checks in f16 GGML_OP_SET_ROWS Stanisław Szymczyk 2026-07-06 11:34:09 +02:00
  • 86961efd56 vulkan: fix 32-bit integer overflow in CEIL_DIV (#25245) b9884 hokanosekai 2026-07-06 10:35:57 +02:00
  • d80e878501 ui: restore Ctrl+B sidebar toggle shortcut (#25307) Pascal 2026-07-06 10:30:07 +02:00
  • 48719618e8 scripts : use HF_TOKEN when downloading UI assets (#25280) b9882 Adrien Gallouët 2026-07-06 09:53:35 +02:00
  • 506f45c26f ggml : add support for CPU f16->f16 GGML_OP_SET_ROWS Stanisław Szymczyk 2026-07-06 08:46:40 +02:00
  • d06ddd3589 ggml-hip: enable -ffast-math for HIP builds (#23862) b9881 a-huk 2026-07-06 09:02:26 +02:00
  • 898b08854d ui: fake 200 for proxy DELETE req (#25298) Xuan-Son Nguyen 2026-07-06 08:41:39 +02:00
  • 72874f559c ggml-cuda: optimize conv_transpose_1d indexing (#25310) b9879 adavyas 2026-07-05 20:49:06 -07:00
  • 2da6686176 Fix stale tensor-split params for draft models (#24814) b9878 Al G 2026-07-05 19:39:36 +01:00
  • 3e5036fbfb abort if we see a multi buffer (#25276) b9877 Eve 2026-07-05 18:38:47 +00:00
  • 4b2a0cdee1 ggml : fix tensor-parallel + -ncmoe crash on MoE models (#25028) b9876 liminfei-amd 2026-07-06 01:56:11 +08:00
  • 7a63fdede1 ggml: Update VMM Pool allocation ggml-cuda.cu - Turing P2P access fix (fixes #24489) (#24491) Vexxie 2026-07-05 18:10:09 +01:00
  • 78d2f52468 cuda : concat implementation for quantized types (#25303) b9874 fairydreaming 2026-07-05 17:26:24 +02:00
  • 124d3d9428 ggml: asynchronous scheduler memory copies jg/ggml-backend-async-copy Johannes Gäßler 2026-07-04 09:49:45 +02:00
  • a4107133a6 llama : add guard for K/V rotation input when buffer is unallocated (#25215) b9873 liminfei-amd 2026-07-05 04:37:38 +08:00
  • 665892536d ui: add sync blocks so display/behavior settings can be set via --ui-config-file (#25132) Pascal 2026-07-04 16:12:27 +02:00
  • ef2d770117 ggml : fix broken CPU concat implementation for quantized types (#25247) b9871 fairydreaming 2026-07-04 13:37:37 +02:00
  • 2d973636e2 chat: trim messages sent to StepFun parser (fixes long reasoning loops) (#25238) b9870 Piotr Wilkin (ilintar) 2026-07-03 23:12:11 +02:00
  • d4cff114c0 ui: Improve performance when streaming (#25225) Nick Towle 2026-07-03 10:03:51 -07:00
  • ff9e7673cf musaDeviceReset xsn/ggml_backend_dev_reset Xuan Son Nguyen 2026-07-03 18:57:48 +02:00
  • 74dc1670fe nits Xuan Son Nguyen 2026-07-03 17:47:39 +02:00
  • 8fb4063720 use CUDA_CHECK Xuan Son Nguyen 2026-07-03 17:47:31 +02:00
  • 602ceaca51 unimpl for other backend Xuan Son Nguyen 2026-07-03 17:37:40 +02:00
  • 2acfc61fe3 add hipDeviceReset() Xuan Son Nguyen 2026-07-03 17:35:25 +02:00
  • f113e02d5a ui: strip path and weight extension from model id in single model mode (#25137) Pascal 2026-07-03 17:32:48 +02:00
  • 893b45a626 add ggml_backend_dev_reset() for sleep mode Xuan Son Nguyen 2026-07-03 17:25:58 +02:00
  • 697372dddb wip xsn/server_sleep_free_backend_2 Xuan Son Nguyen 2026-07-03 17:15:50 +02:00
  • 7f88a1b9f9 move it to ggml_backend_reg_i Xuan Son Nguyen 2026-07-03 16:35:19 +02:00
  • 4d990e1113 free cuda device in ggml_backend_unload Xuan Son Nguyen 2026-07-03 16:26:09 +02:00
  • 152d337fad spec: support spec-draft-p-min in DFlash (#25246) b9867 Ruixiang Wang 2026-07-03 15:40:06 +02:00
  • 75a48a9055 cuda: enable topk-moe fusion for 288 experts (#25267) b9866 Piotr Wilkin (ilintar) 2026-07-03 15:36:55 +02:00
  • 132e86bcfb server: also free backen on sleep Xuan Son Nguyen 2026-07-03 13:21:34 +02:00
  • 067de93718 ui: align persisted config with strict server schema and enable thinking by default (#25242) Pascal 2026-07-03 13:14:52 +02:00
  • b5315e16e0 server + ui: ping silent SSE streams every 1s and kick only after 3s so slow prefill never drops healthy connections (#25241) b9864 Pascal 2026-07-03 12:47:04 +02:00
  • 94875285e4 ui: Add MCP Servers Opt-In for first time visitors (#25239) Aleksander Grygier 2026-07-03 12:16:29 +02:00
  • 5a460dea9f Remove redundant CUDA copies after gated_delta_net. (#23940) b9862 Gaurav Garg 2026-07-03 14:36:29 +05:30
  • c8ae9a750c vendor : update cpp-httplib to 0.49.0 (#25218) b9861 Alessandro de Oliveira Faria (A.K.A.CABELO) 2026-07-03 05:26:54 -03:00
  • fdb1db877c llama : add llama_model_ftype_name() (#25134) b9860 Adrien Gallouët 2026-07-02 17:26:47 +02:00
  • 4fc4ec5541 opencl: allow loading precompiled binary kernels from library (#23042) b9859 lhez 2026-07-01 10:29:22 -07:00
  • a6647b1a32 common : use hf primary split as model path (#25194) b9858 Adrien Gallouët 2026-07-01 18:33:00 +02:00
  • 13e673863b hexagon: flash attention rework (optimizations, accuracy improvements, etc) (#25085) b9857 Max Krasnyansky 2026-07-01 06:59:19 -07:00
  • b820cc8e6f CUDA: consistent use of __restrict__ + PDL for FA (#25185) b9856 Johannes Gäßler 2026-07-01 10:55:14 +02:00
  • 6dbc1174b8 ggml-cpu: add AVX2 optimization for nvfp4 dot product and use UE4M3 LUT (#23961) b9855 ragz4125 2026-07-01 13:01:20 +05:30
  • 9d88e7cedd ui Prevent tool messages from incorrectly appending to other conversations (#25177) Aleksander Grygier 2026-07-01 09:25:18 +02:00
  • 7af4279f45 ui: Remove PWA navigate fallback to prevent caching API endpoint requests (#25174) b9853 Aleksander Grygier 2026-07-01 07:32:55 +02:00
  • fd1a05791d opencl: initial q1_0 support (#25160) b9852 lhez 2026-06-30 21:43:20 -07:00
  • 0eca4d490e cuda : prevent integer truncation and overflow errors when using KQ mask strides in flash_attn_mask_to_KV_max kernel (#24945) b9851 fairydreaming 2026-06-30 20:47:05 +02:00
  • 4f31eedb0c model : register t_layer_inp for qwen3next (#25141) b9850 Jürgen Schmied 2026-06-30 17:57:14 +02:00