Commit Graph

  • f7b1116af1 update release requirements (#11897) b4734 Eve 2025-02-17 11:20:23 +00:00
  • c4d29baf32 server : fix divide-by-zero in metrics reporting (#11915) b4733 Antoine Viallon 2025-02-17 11:25:12 +01:00
  • 2eea03d86a vulkan: implement several ops relevant for ggml_opt (#11769) b4732 Rémy O 2025-02-17 07:55:57 +01:00
  • 0f2bbe6564 server : bump httplib to 0.19.0 (#11908) b4731 Xuan-Son Nguyen 2025-02-16 18:11:22 +01:00
  • aed4a8e980 fix server Xuan Son Nguyen 2025-02-16 11:36:50 +01:00
  • fe163d5bf3 common : Fix a typo in help (#11899) b4730 standby24x7 2025-02-16 18:51:13 +09:00
  • 818a340ea8 ci : fix (again) arm64 build fails (#11895) Xuan-Son Nguyen 2025-02-16 10:36:39 +01:00
  • bf42a23d0a vulkan: support multi/vision rope, and noncontiguous rope (#11902) b4728 Jeff Bolz 2025-02-16 01:52:23 -06:00
  • c2ea16f260 metal : fix the crash caused by the lack of residency set support on Intel Macs. (#11904) b4727 Hale Chan 2025-02-16 14:50:26 +08:00
  • 85ef80cbe9 server : use llama_batch_ext Xuan Son Nguyen 2025-02-16 00:06:48 +01:00
  • 17d3658b5f move to llama_batch_ext Xuan Son Nguyen 2025-02-16 00:02:53 +01:00
  • 6dde178248 scripts: fix compare-llama-bench commit hash logic (#11891) Johannes Gäßler 2025-02-15 20:23:22 +01:00
  • fc10c38ded examples: fix typo in imatrix/README.md (#11884) 708-145 2025-02-15 20:03:30 +01:00
  • 22885105a6 metal : optimize dequant q6_K kernel (#11892) b4724 Adrian Kretz 2025-02-15 19:39:20 +01:00
  • c2cd24fbfd readme : add notice about new package registry (#11890) Georgi Gerganov 2025-02-15 20:29:56 +02:00
  • 68ff663a04 repo : update links to new url (#11886) b4722 Georgi Gerganov 2025-02-15 16:40:57 +02:00
  • 8654805027 docker : publish to both ggerganov and ggml-org xsn/ci_legacy_gg Xuan Son Nguyen 2025-02-15 15:18:04 +01:00
  • f355229692 server: fix type promotion typo causing crashes w/ --jinja w/o tools (#11880) b4721 Olivier Chafik 2025-02-15 10:11:36 +00:00
  • fc1b0d0936 vulkan: initial support for IQ1_S and IQ1_M quantizations (#11528) b4720 Rémy O 2025-02-15 09:01:40 +01:00
  • 89daa2564f llguidance build fixes for Windows (#11664) b4719 Michał Moskal 2025-02-14 12:46:08 -08:00
  • 300907b211 opencl: Fix rope and softmax (#11833) b4718 lhez 2025-02-14 11:12:23 -08:00
  • f2e59a8eb9 rework, targeting llama-server Xuan Son Nguyen 2025-02-14 18:16:49 +01:00
  • 1d801d27b9 graph : update attn/kv_self names Georgi Gerganov 2025-02-14 17:22:55 +02:00
  • 828064564c context : move common inputs to base class Georgi Gerganov 2025-02-14 16:48:21 +02:00
  • 94b87f87b5 cuda : add ampere to the list of default architectures (#11870) b4717 Diego Devesa 2025-02-14 15:33:52 +01:00
  • d5e8e1a2ba context : remove batch_manager Georgi Gerganov 2025-02-14 16:10:55 +02:00
  • dbc2ec59b5 docker : drop to CUDA 12.4 (#11869) b4716 Georgi Gerganov 2025-02-14 14:48:40 +02:00
  • 3d68f034da llama : add completion for --chat-template-file (#11860) Daniel Bevenius 2025-02-14 11:16:56 +01:00
  • 38e32eb6a0 ggml: optimize some vec dot functions for LoongArch ASX (#11842) b4714 Jinyang He 2025-02-14 16:54:27 +08:00
  • a4f011e8d0 vulkan: linux builds + small subgroup size fixes (#11767) b4713 Eve 2025-02-14 02:59:40 +00:00
  • a7b8ce2260 llama-bench : fix unexpected global variable initialize sequence issue (#11832) b4712 theraininsky 2025-02-14 09:13:43 +08:00
  • 4ed4fe75ed first proposal for private llama_batch Xuan Son Nguyen 2025-02-14 00:48:12 +01:00
  • 04045bb842 readme : minor Georgi Gerganov 2025-02-14 00:16:56 +02:00
  • 8a8c4ceb60 llamafile: use member variable instead of constant for iq4nlt (#11780) b4710 Jeffrey Morgan 2025-02-13 09:05:04 -08:00
  • c1f958c038 server : (docs) Update wrong tool calling example (#11809) Reza Rahemtola 2025-02-13 17:22:44 +01:00
  • 131743ff4f context : abstract constructor and init Georgi Gerganov 2025-02-13 17:13:42 +02:00
  • ed3cb55abe context : abstract input Georgi Gerganov 2025-02-13 15:53:15 +02:00
  • c48f630d1c llama : add --completion-bash option (#11846) b4708 Daniel Bevenius 2025-02-13 14:46:59 +01:00
  • 107d1e2c32 context : move output functionality to base class Georgi Gerganov 2025-02-13 15:42:14 +02:00
  • bd6e55bfd3 musa: bump MUSA SDK version to rc3.1.1 (#11822) b4707 R0CKSTAR 2025-02-13 20:28:18 +08:00
  • e08f38df69 context : minor cleanup Georgi Gerganov 2025-02-13 12:50:53 +02:00
  • f7c7757bab context : abstract state read/write Georgi Gerganov 2025-02-13 12:37:28 +02:00
  • 3a504d9a0b llama : introduce llama_io interfaces Georgi Gerganov 2025-02-13 12:18:44 +02:00
  • c7f460ab88 server: fix tool-call of DeepSeek R1 Qwen, return reasoning_content (Command 7RB & DeepSeek R1) unless --reasoning-format none (#11607) b4706 Olivier Chafik 2025-02-13 10:05:16 +00:00
  • 27e8a23300 sampling: add Top-nσ sampler (#11223) b4705 Vinesh Janarthanan 2025-02-13 00:45:57 -06:00
  • e4376270d9 llama.cpp: fix warning message (#11839) b4704 Oleksandr Kuvshynov 2025-02-13 01:25:34 -05:00
  • 3e69319772 llama : update llama_decode_internal ref [no ci] (#11840) Daniel Bevenius 2025-02-13 07:07:51 +01:00
  • a394039db0 ggml-cpu : add chunking support to mul_mat_id (#11666) b4702 Diego Devesa 2025-02-13 01:02:38 +01:00
  • be3bbd6215 ggml : x2 speed for WASM by optimizing SIMD (#11453) Xuan-Son Nguyen 2025-02-13 00:33:45 +01:00
  • 31afcbee0e server : (webui) Give copy button back to all message bubbles (#11814) Woof Dog 2025-02-12 22:47:11 +00:00
  • 5c4284d57b HIP: Remove GCN from list of devices that avoid MMQ (#11831) b4699 uvos 2025-02-12 22:25:28 +01:00
  • bfd11a2344 Fix: Compile failure due to Microsoft STL breaking change (#11836) b4698 JC 2025-02-12 20:36:11 +00:00
  • 0fb77f821f sync : ggml Georgi Gerganov 2025-02-12 21:46:02 +02:00
  • f30aca84b2 Revert "HIP: Switch to std::vector in rocblas version check (#11820)" revert-11820-vers_fix uvos 2025-02-12 19:22:04 +01:00
  • e598697d63 HIP: Switch to std::vector in rocblas version check (#11820) b4696 uvos 2025-02-12 17:25:03 +01:00
  • fbe6a07256 context : rename to llama_context_kv_self Georgi Gerganov 2025-02-12 17:16:44 +02:00
  • 6ee86e5e0f graph : restore ubatch in build_cb Georgi Gerganov 2025-02-12 16:29:15 +02:00
  • fef0cbeadf cleanup: fix compile warnings associated with gnu_printf (#11811) b4695 bandoti 2025-02-12 10:06:53 -04:00
  • 748ee9fe93 ggml : fix multi-threaded clamp_f32 (#11824) b4694 Richard 2025-02-12 13:57:33 +00:00
  • f63aeecce6 llama : models now build their graphs using llama_graph_i Georgi Gerganov 2025-02-12 15:08:40 +02:00
  • 198b1ec611 ggml-cpu: Fix duplicate MATMUL_INT8 (#11817) Weizhao Ouyang 2025-02-12 20:22:58 +08:00
  • c3d6af7cd2 CUDA: fix CUDART_VERSION checks (#11821) b4692 Johannes Gäßler 2025-02-12 13:16:39 +01:00
  • 0ab50f1bbb context : prepare llama_model graph build Georgi Gerganov 2025-02-12 13:59:43 +02:00
  • e633dc171a context : introduce llama_graph_i Georgi Gerganov 2025-02-12 13:48:52 +02:00
  • 5eae8e5183 context : move build_rope_factors to base class Georgi Gerganov 2025-02-12 13:32:02 +02:00
  • d146a14f77 context : minor naming fix Georgi Gerganov 2025-02-12 12:41:36 +02:00
  • 8da7f612b7 context : improve llama_context encapsulation Georgi Gerganov 2025-02-12 12:11:30 +02:00
  • b52b79b048 context : move encode/decode to llama-context.cpp Georgi Gerganov 2025-02-12 11:23:38 +02:00
  • 369be5598a llama : fix typo in llama-grammar.h [no ci] (#11816) Daniel Bevenius 2025-02-12 08:40:01 +01:00
  • 4078c77f98 docs: add OpenCL (#11697) lhez 2025-02-11 14:04:13 -08:00
  • 02ef4be975 context : initial abstraction Georgi Gerganov 2025-02-11 11:25:18 +02:00
  • 90e4dba461 Fix #11802: Compile bug - RegQueryValueExA changed to RegQueryValueEx (#11803) b4689 Sheldon Robinson 2025-02-11 10:55:45 -05:00
  • a18f481f99 server : use common_token_to_piece instead of common_detokenize (#11740) b4688 Daniel Bevenius 2025-02-11 14:06:45 +01:00
  • b9ab0a4d0b CUDA: use arch list for compatibility check (#11775) Johannes Gäßler 2025-02-11 00:17:22 +01:00
  • 7b891bdc86 fix: typos in documentation files (#11791) b4686 Maxim Evtush 2025-02-10 23:21:31 +01:00
  • 81732619fd docs: utilize the forward slash (/) as the path separator for Unix-like systems (#11770) jason_w 2025-02-11 06:17:48 +08:00
  • 507f9174fe server : (webui) introduce conversation branching + idb storage (#11792) Xuan-Son Nguyen 2025-02-10 21:23:17 +01:00
  • 19b392d58d llama-mmap: fix missing include (#11796) b4683 Wilken Gottwalt 2025-02-10 19:58:18 +01:00
  • 0893e0114e server : correct signal handler (#11795) b4682 Xuan-Son Nguyen 2025-02-10 18:03:28 +01:00
  • 2cd8a903c8 context : make output functions members Georgi Gerganov 2025-02-10 17:01:27 +02:00
  • d1d8d53008 bman : remove ubatch member Georgi Gerganov 2025-02-10 16:50:14 +02:00
  • ef358ee78f context : add decode/encode Georgi Gerganov 2025-02-10 16:11:17 +02:00
  • 879ba82777 server : increase context size for the tests Georgi Gerganov 2025-02-10 15:00:02 +02:00
  • f9971ef2e1 llama : dedup reserve code Georgi Gerganov 2025-02-10 14:59:51 +02:00
  • 972f91c7d7 Merge branch 'master' into gg/llama-kv-cache Georgi Gerganov 2025-02-10 14:45:54 +02:00
  • d7b31a9d84 sync: minja (a72057e519) (#11774) b4681 Olivier Chafik 2025-02-10 09:34:09 +00:00
  • 9ac3457b39 Update README.md [no ci] (#11781) pascal-lc 2025-02-10 16:05:57 +08:00
  • c2a67efe38 vulkan: Make Vulkan optional at runtime (#11493). (#11494) b4679 Danny Milosavljevic 2025-02-10 07:17:21 +01:00
  • b044a0fe3c vulkan: add environment variable GGML_VK_PREFER_HOST_MEMORY to avoid VRAM allocation (#11592) b4678 Wagner Bruna 2025-02-10 03:08:22 -03:00
  • 1be357d990 Merge branch 'master' into compilade/imatrix-batched-chunks Francis Couture-Harpin 2025-02-09 12:06:24 -05:00
  • db502ddd0e Merge branch 'master' into compilade/imatrix-batched-chunks Francis Couture-Harpin 2025-02-09 12:06:15 -05:00
  • 19d3c8293b There's a better way of clearing lines (#11756) b4677 Eric Curtin 2025-02-09 10:34:49 +00:00
  • 98f6b0fd1e vulkan: account for lookup tables when checking shared memory size (#11502) b4676 Jeff Bolz 2025-02-09 01:43:51 -06:00
  • 55ac8c7791 server : (webui) revamp Settings dialog, add Pyodide interpreter (#11759) b4675 Xuan-Son Nguyen 2025-02-08 21:54:50 +01:00
  • e6e6583199 server : (webui) increase edit textarea size (#11763) Woof Dog 2025-02-08 19:09:55 +00:00
  • aaa5505307 server : minor log updates (#11760) Georgi Gerganov 2025-02-08 18:08:43 +02:00
  • bdcf8b6a56 cont : fix mmap flag print (#11699) Georgi Gerganov 2025-02-08 16:49:38 +02:00
  • 4d3465c5ae ggml: Fix data race in ggml threadpool (#11736) b4671 Karol Kontny 2025-02-08 15:30:53 +01:00
  • d86e23101e server : minor log updates gg/server-logs Georgi Gerganov 2025-02-08 16:23:37 +02:00
  • d80be897ac CUDA: fix min. version for movmatrix (#11751) Johannes Gäßler 2025-02-08 10:46:07 +01:00