Commit Graph

  • 074e42ab31 convert : converting mmproj for Qwen2/2.5VL from convert_hf_to_gguf (#13209) Xuan-Son Nguyen 2025-05-02 17:17:15 +02:00
  • c642bc014c kv-cache : separate recurrent vs non-recurrent impl (#12799) Georgi Gerganov 2025-05-02 17:48:36 +03:00
  • cb06a3c363 llama : orion rope type is neox (#13261) b5261 Sigbjørn Skjæret 2025-05-02 12:44:24 +02:00
  • 626083faf7 llama : plamo rope type is neox (#13260) b5260 Sigbjørn Skjæret 2025-05-02 12:40:56 +02:00
  • 2af6880178 llama-chat : reset glmedge chat template (#13253) b5259 piDack 2025-05-02 17:06:09 +08:00
  • e84773ab60 mtmd-cli : fix out_of_range when input image path is empty (#13244) b5258 Shakil Ahmed 2025-05-02 14:20:27 +06:00
  • fab647e884 server : add cache reuse card link to help (#13230) b5257 Georgi Gerganov 2025-05-02 09:48:31 +03:00
  • dcf886007d convert : explicitly disable trust_remote_code for AutoConfig (#13246) Xuan-Son Nguyen 2025-05-02 08:45:10 +02:00
  • 94c3d53043 kv-cache : remove const_cast when setting inputs for s_copy Francis Couture-Harpin 2025-05-01 22:18:57 -04:00
  • 791998b42d metal : single-user mamba2 inference works Francis Couture-Harpin 2025-05-01 21:27:12 -04:00
  • 6def5cd729 metal : add missing args for nb references in ssm_scan_f32_group Francis Couture-Harpin 2025-05-01 19:10:20 -04:00
  • cf4f0a4123 metal : fix confusion between ; and , Francis Couture-Harpin 2025-05-01 18:55:34 -04:00
  • d24d592808 ci: fix cross-compile sync issues (#12804) b5255 bandoti 2025-05-01 19:06:39 -03:00
  • 35d06fac5a Merge branch 'master' into compilade/mamba2 Francis Couture-Harpin 2025-05-01 17:43:53 -04:00
  • 8efbdadc61 rpc : avoid uninitialized memory in serialize_tensor (#13210) b5254 Justin Santa Barbara 2025-05-01 17:32:11 -04:00
  • f057808ffa ggml: Don't assert fail when tensor data changes (#13222) b5253 Jesse Gross 2025-05-01 13:46:10 -07:00
  • d7a14c42a1 build : fix build info on windows (#13239) b5252 Diego Devesa 2025-05-01 21:48:08 +02:00
  • b6e4ff69b8 clip : (minicpmv) Re-enable upscaling of images smaller than the CLIP image size (#13237) Loïc Carrère 2025-05-01 21:32:21 +02:00
  • e0f572c846 llama-chat : update GLM4 chat template (#13238) b5250 matteo 2025-05-01 21:16:38 +02:00
  • 79f26e9e12 vulkan: Add bfloat16 support (#12554) b5249 Jeff Bolz 2025-05-01 13:49:39 -05:00
  • fc727bcdd5 vulkan: Handle src1 batch dimension in non-contiguous mat-vec-mul shader (#13191) b5248 Jeff Bolz 2025-05-01 13:19:31 -05:00
  • b0ecbd434b test: non-cont. b in test-backend-ops -o MUL_MAT (#13187) Johannes Gäßler 2025-05-01 20:18:56 +02:00
  • b1dd4d08e8 sync : ggml b5246 Georgi Gerganov 2025-05-01 17:07:13 +03:00
  • 99881f77d8 whisper : add check that target name exists (whisper/3103) Daniel Bevenius 2025-05-01 10:05:24 +02:00
  • b5769d92b4 ggml : suppress Windows compiler warnings (whisper/3075) Daniel Bevenius 2025-04-29 15:47:55 +02:00
  • 8936784f7a mtmd : add **vision** support for Mistral Small 3.1 (#13231) b5243 Xuan-Son Nguyen 2025-05-01 17:05:42 +02:00
  • 13c9a3319b arg : remove CURLINFO_EFFECTIVE_METHOD (#13228) b5242 Xuan-Son Nguyen 2025-05-01 10:23:25 +02:00
  • a70183eb00 llama-model : fix the reported size class for nomic-embed-text-v2-moe (#13223) b5241 Jared Van Bortel 2025-05-01 03:09:41 -04:00
  • 8d33d740c3 sync : ggml Georgi Gerganov 2025-05-01 09:59:02 +03:00
  • 65202d2985 sync : ggml sync-ggml-25-05-01 Georgi Gerganov 2025-05-01 09:59:02 +03:00
  • 4254bb4951 ggml : fix ggml_gallocr_ptr type (ggml/1205) b5239 Diego Devesa 2025-04-30 15:20:40 +02:00
  • 9998540149 cuda : fix unused variable compile warning (whisper/0) Georgi Gerganov 2025-04-24 18:59:06 +03:00
  • db1ff5b63a ggml : fix ggml_gallocr_ptr type (ggml/1205) Diego Devesa 2025-04-30 15:20:40 +02:00
  • 610df4cc3b cuda : fix unused variable compile warning (whisper/0) Georgi Gerganov 2025-04-24 18:59:06 +03:00
  • e1e8e0991f CUDA: batched+noncont MMQ, refactor bs>1 MoE code (#13199) b5237 Johannes Gäßler 2025-04-30 23:12:59 +02:00
  • 6f67cf1f48 arg : -hf do not fail if url mismatch (#13219) b5236 Xuan-Son Nguyen 2025-04-30 22:29:15 +02:00
  • 16a457facd fix typo: n_ctx_pre_seq -> n_ctx_per_seq (#13221) b5235 ddh0 2025-04-30 15:28:43 -05:00
  • 3e168bede4 convert : improve model arch handling (#13122) Xuan-Son Nguyen 2025-04-30 16:56:24 +02:00
  • ceda28ef8e llava : remove duplicate include (#13207) b5233 Tatsuya Tanaka 2025-04-30 22:25:20 +09:00
  • 3b127c7385 common : add -jf / --json-schema-file flag (#12011) b5232 Olivier Chafik 2025-04-30 13:52:35 +01:00
  • e5007a5edf vulkan: use uint array index to avoid glslang bug (#13193) b5231 Jeff Bolz 2025-04-30 07:38:37 -05:00
  • 416313773b ggml : fix ppc64le build (#13176) b5230 shalinib-ibm 2025-04-30 16:47:08 +05:30
  • 07c2e2f76c convert : correct typo image_mean --> image_std (#13208) Xuan-Son Nguyen 2025-04-30 13:06:15 +02:00
  • 44cd8d91ff feat(ggml-cpu): enable z17 compile (#13182) b5228 Aaron Teo 2025-04-30 17:47:35 +08:00
  • 5933e6fdc9 arg : allow using -hf offline (#13202) Xuan-Son Nguyen 2025-04-30 10:46:32 +02:00
  • da84c04d8f docker : do not build tests (#13204) b5226 Xuan-Son Nguyen 2025-04-30 10:44:07 +02:00
  • a0f7016d17 rpc : fix cache directory initialization (#13188) b5225 xiaofei 2025-04-30 14:29:22 +08:00
  • 19e899ce21 scripts: n_depth for compare-llama-bench [no ci] (#13201) Johannes Gäßler 2025-04-29 23:32:04 +02:00
  • e2e1ddb93a server : Prefilling assistant message in openai compatible API (#13174) b5223 matteo 2025-04-29 20:33:10 +02:00
  • d9d398f84f sampling : when top-k <= 0 -> noop (#13173) b5222 Georgi Gerganov 2025-04-29 20:22:57 +03:00
  • 5a63980117 llama-bench: fixed size of fields to correctly map to values (#13183) b5221 Alberto Cabrera Pérez 2025-04-29 16:24:36 +01:00
  • cdf76586b2 CUDA: fix non-cont. inputs for batched mat mul (#13155) b5220 Johannes Gäßler 2025-04-29 16:00:27 +02:00
  • 7d3af70b08 llama : llm_type order by size (#13177) b5219 Sigbjørn Skjæret 2025-04-29 13:25:53 +02:00
  • 00e3e5a194 mtmd : add qwen2vl and qwen2.5vl (#13141) b5218 Xuan-Son Nguyen 2025-04-29 11:47:04 +02:00
  • e98b3692be llama : set qwen3 model type sizes (#13175) b5217 Sigbjørn Skjæret 2025-04-29 11:00:31 +02:00
  • b6ce7430b7 llama-graph : fix text position for mrope (#13159) b5216 Xuan-Son Nguyen 2025-04-29 08:45:49 +02:00
  • 5f5e39e1ba model : Nomic Embed Text V2 with Mixture-of-Experts (MoE) architecture (#12466) b5215 AT 2025-04-28 15:52:15 -04:00
  • eaea325324 clip : fix model size display (#13153) b5214 Xuan-Son Nguyen 2025-04-28 21:23:19 +02:00
  • 43ddab6eee fix(rpc): Improve input validation and error handling (#13069) b5213 Ville Vesilehto 2025-04-28 21:00:20 +03:00
  • 1831f538f7 llama-bench: add -d depth arg (#13096) b5212 Vishal Agarwal 2025-04-28 20:20:39 +05:30
  • 4e87962e34 mtmd : fix glm-edge redundant token count (#13139) b5211 Xuan-Son Nguyen 2025-04-28 16:12:56 +02:00
  • fb0471d175 context : do not clear output buffer on reserve (#13152) b5210 pockers21 2025-04-28 06:45:40 -07:00
  • d2b2031e5f llama : (mrope) allow using normal 1D position for text token (#13138) b5209 Xuan-Son Nguyen 2025-04-28 14:20:56 +02:00
  • 5fa9e63be8 clip : refactor set input for cgraph + fix qwen2.5vl input (#13136) b5208 Xuan-Son Nguyen 2025-04-28 12:18:59 +02:00
  • a4c340f974 SYCL: Add all missing unary kernels (#13074) b5207 Akarshan Biswas 2025-04-28 15:03:25 +05:30
  • d0a417f3c7 readme : update hot topics (#13150) Georgi Gerganov 2025-04-28 12:10:18 +03:00
  • 43f2b07193 common : fix noreturn compile warning (#13151) b5205 Georgi Gerganov 2025-04-28 11:57:19 +03:00
  • e5d6c2554e llama-chat : fix typo GML --> GLM (#13143) b5204 Xuan-Son Nguyen 2025-04-28 10:11:58 +02:00
  • b710758323 readme : update hot topics gg/survey-nvidia Georgi Gerganov 2025-04-28 10:46:43 +03:00
  • f0dd6a1926 musa: fix typo in cc control (#13144) R0CKSTAR 2025-04-28 15:33:28 +08:00
  • 69699be48a CUDA: fix q_nope_absorbed prec for DS 2 Lite f16 (#13137) b5202 Johannes Gäßler 2025-04-28 09:29:26 +02:00
  • 85f36e5e71 arg : fix unused variable (#13142) b5201 Xuan-Son Nguyen 2025-04-28 07:16:59 +02:00
  • c0a97b762e llama-bench : Add --override-tensors arg (#12922) b5200 4onen 2025-04-27 14:48:26 -07:00
  • ced44be342 llama-chat : fix wrong template in GLM4-0414 (#13140) b5199 matteo 2025-04-27 21:57:32 +02:00
  • e291450b76 musa: fix build warning (#13129) b5198 R0CKSTAR 2025-04-27 19:22:49 +08:00
  • 59e991c23c Fixes Qwen2.5VL segfault during inference with https://github.com/ggml-org/llama.cpp/pull/12402 as has_qwen2vl_merger migration was incomplete (#13133) b5197 LostRuins Concedo 2025-04-27 18:43:37 +08:00
  • 37ae6a281a Fixes Qwen2.5VL segfault during inference with https://github.com/ggml-org/llama.cpp/pull/12402 as has_qwen2vl_merger migration was incomplete cedo/fix-q25vl Concedo 2025-04-27 18:36:57 +08:00
  • ca2bb89eac clip : Add Qwen2.5VL support (#12402) b5196 HimariO 2025-04-27 16:10:34 +08:00
  • 2d451c8059 common : add common_remote_get_content (#13123) b5195 Xuan-Son Nguyen 2025-04-26 22:58:12 +02:00
  • 4753791e70 clip : improve projector naming (#13118) b5194 Xuan-Son Nguyen 2025-04-26 22:39:47 +02:00
  • 77d5e9a76a ggml: move fp16/bf16 conversion optimizations to CPU backend + export conversion APIs (#13107) b5193 SXX 2025-04-26 22:05:31 +08:00
  • d5fe4e81bd grammar : handle maxItems == 0 in JSON schema (#13117) b5192 frob 2025-04-26 10:10:20 +02:00
  • 295354ea68 llama : fix K-shift with quantized K and BLAS backend (#13113) b5191 Diego Devesa 2025-04-25 19:40:11 +02:00
  • ed68474f76 wip gg/model-cards Georgi Gerganov 2025-04-25 19:07:09 +03:00
  • 558a764713 Force FP32 compute in GLM4 FFN Down (#13101) b5190 City 2025-04-25 14:38:34 +02:00
  • edb18b6e8f clip : fix pixtral on some GPU backends (#13097) b5189 Xuan-Son Nguyen 2025-04-25 14:31:42 +02:00
  • 514c45608f change the reorder tensor from init to execute OP (#13003) b5188 Neo Zhang Jianyu 2025-04-25 17:37:51 +08:00
  • 553a5c3a9f rpc : do not wait for response when sending RPC_CMD_SET_TENSOR (#12943) b5187 Radoslav Gerganov 2025-04-25 10:08:08 +03:00
  • 13be08daf9 clip : remove boi/eoi embeddings for GLM-edge model (#13081) b5186 Xuan-Son Nguyen 2025-04-24 22:17:04 +02:00
  • 226251ed56 embeddings : fix batch sizes (#13076) b5185 Georgi Gerganov 2025-04-24 22:29:22 +03:00
  • 87616f0680 ggml : fix trailing whitespaces (#0) b5184 Georgi Gerganov 2025-04-24 17:22:27 +03:00
  • 63b4911494 sync : ggml Georgi Gerganov 2025-04-24 16:47:43 +03:00
  • c6e8cc28c1 ggml : Depthwise 2D convolution (ggml/1152) Acly 2025-04-17 14:16:45 +02:00
  • b10d8bfdb1 CUDA: use switch statements in constexpr functions (#13095) b5181 Johannes Gäßler 2025-04-24 15:57:10 +02:00
  • 13b4548877 cmake : do not include ./src as public for libllama (#13062) b5180 Georgi Gerganov 2025-04-24 16:00:10 +03:00
  • 572b3141d3 clang-tidy : disable warning about missing math parenthesis (#13091) Georgi Gerganov 2025-04-24 15:44:05 +03:00
  • 7c727fbe39 arg : add --no-mmproj-offload (#13093) b5178 Xuan-Son Nguyen 2025-04-24 14:04:14 +02:00
  • 80982e815e arg : clean up handling --mmproj with -hf (#13082) b5177 Xuan-Son Nguyen 2025-04-24 12:14:13 +02:00
  • 7604a7d6b8 metal : fix floating-point range of attention scores in FA kernels (#13090) b5176 Georgi Gerganov 2025-04-24 10:38:30 +03:00
  • b3b6d862cf vulkan: matmul gcn tuning (#13016) b5175 Eve 2025-04-24 07:18:33 +00:00