Commit Graph

  • 2c3f8b850a llama : support BailingMoE (Ling) (#12634) b5002 Sigbjørn Skjæret 2025-03-30 22:21:03 +02:00
  • 4663bd353c metal : use constexpr in FA kernels + fix typedef (#12659) b5001 Georgi Gerganov 2025-03-30 22:04:04 +03:00
  • b3de7cac73 llama : add Trillion 7B model support (#12556) Juyoung Suk 2025-03-31 03:38:33 +09:00
  • 7242dd9675 llama-chat : Add Yandex instruct model template support (#12621) b4999 Sergei Vorobyov 2025-03-30 21:12:03 +03:00
  • 492d7f1ff7 musa: fix all warnings, re-enable -DLLAMA_FATAL_WARNINGS=ON in ci and update doc (#12611) b4998 R0CKSTAR 2025-03-30 16:59:38 +08:00
  • d3f1f0acfb sync : ggml b4997 Georgi Gerganov 2025-03-29 15:37:54 +02:00
  • 360dc22c00 cpu : rm unused variable (ggml/1166) Xuan-Son Nguyen 2025-03-29 11:59:56 +01:00
  • a62d7fa7a9 cpu: de-duplicate some of the operators and refactor (ggml/1144) cmdr2 2025-03-29 11:37:13 +05:30
  • e408d4351a ggml : add logging for native build options/vars (whisper/2935) Daniel Bevenius 2025-03-24 09:53:38 +01:00
  • 3891e183c6 examples : command.wasm updates (whisper/2904) Daniel Bevenius 2025-03-20 07:02:18 +01:00
  • af6ae1efb2 llama : fix non-causal mask for gemma 3 (#12615) b4992 Xuan-Son Nguyen 2025-03-30 00:07:37 +01:00
  • 0bb2919335 llama : change cpu_buft_list order: ACCEL -> GPU host -> CPU extra -> CPU (#12632) b4991 Djip007 2025-03-29 14:07:37 +01:00
  • a69f846351 cmake : fix ccache conflict (#12522) b4990 Jay 2025-03-29 18:04:58 +08:00
  • d07a0d7a79 CANN : remove clang-format in ggml-cann (#12607) hipudding 2025-03-29 18:03:28 +08:00
  • 3714c3ee1a llama : fix incorrect Qwen2Moe ffn_moe_out graph callback (#12631) b4988 Sigbjørn Skjæret 2025-03-28 22:13:02 +01:00
  • b4ae50810e metal : improve FA + improve MoE (#12612) b4987 Georgi Gerganov 2025-03-28 20:21:59 +02:00
  • b86f600723 vulkan: fix coopmat shader generation when cross-compiling (#12272) b4986 Icenowy Zheng 2025-03-29 01:51:06 +08:00
  • dd373dd3bf llama: fix error on bad grammar (#12628) b4985 Johannes Gäßler 2025-03-28 18:08:52 +01:00
  • 5d01670266 server : include speculative decoding stats when timings_per_token is enabled (#12603) b4984 Benson Wong 2025-03-28 01:05:44 -07:00
  • ef03229ff4 rpc : update README for cache usage (#12620) Radoslav Gerganov 2025-03-28 09:44:13 +02:00
  • 13731766db llamafile : ppc64le GEMV forwarding for FP32. (#12594) b4982 amritahs-ibm 2025-03-28 13:13:22 +05:30
  • c875e03f96 rpc : update README for cache usage rpc-hash-readme Radoslav Gerganov 2025-03-28 09:15:09 +02:00
  • ab6ab8f809 rpc : send hash when tensor data is above some fixed threshold (#12496) b4981 Radoslav Gerganov 2025-03-28 08:18:04 +02:00
  • 2099a9d5db server : Support listening on a unix socket (#12613) b4980 Piotr 2025-03-27 23:41:04 +01:00
  • 2969019837 media : add SVG logo [no ci] (#12616) Georgi Gerganov 2025-03-27 23:09:05 +02:00
  • efe0222130 media : add SVG logo [no ci] gg/media-add-svg-logo Georgi Gerganov 2025-03-27 23:07:46 +02:00
  • 5dec47dcd4 opencl: add multi and vision rope, gelu_quick and im2col (#12600) b4978 lhez 2025-03-27 08:08:08 -07:00
  • f125b8dccf llama : add PLM GGUF Conversion & Inference Support (#12457) b4977 Si1w 2025-03-27 10:49:15 +00:00
  • 953c2a62cf model : restore support for T5Encoder (#12590) b4976 HighDoping 2025-03-27 18:43:33 +08:00
  • d5c6309d91 convert : Support Qwen2_5_VLForConditionalGeneration (#12595) Csaba Kecskemeti 2025-03-27 03:11:23 -07:00
  • 029c693fdc sync : ggml b4974 Georgi Gerganov 2025-03-27 09:36:13 +02:00
  • 771d84371c scripts : update sync + fix cmake merge Georgi Gerganov 2025-03-27 09:22:30 +02:00
  • df0665a483 sync : ggml b4972 Georgi Gerganov 2025-03-27 09:01:21 +02:00
  • 0306aad1ca cmake : sync/merge PowerPC build commands (#0) Georgi Gerganov 2025-03-27 09:00:57 +02:00
  • c7b43ab608 llamafile : ppc64le MMA implementation for Q4_0. (#12489) b4970 amritahs-ibm 2025-03-27 12:21:47 +05:30
  • 24feaec057 ggml : riscv: add 128-bit RVV support (#12530) b4969 xctan 2025-03-27 14:38:34 +08:00
  • f28bc4c286 llama : make loras compatible with repacking (#12593) Georgi Gerganov 2025-03-27 08:24:10 +02:00
  • f17a3bb4e8 SYCL: implement memset ggml backend buffer interface (#12580) b4967 Akarshan Biswas 2025-03-27 07:16:00 +05:30
  • bd40678df7 HIP: Add support for RDNA4 targets (#12372) b4966 Slobodan Josic 2025-03-26 23:46:30 +01:00
  • b3298fa47a metal : refactor mat-vec code (#12569) Georgi Gerganov 2025-03-26 21:38:38 +02:00
  • 70b063a550 metal : reduce register pressure gg/metal-refactor-mv-2 Georgi Gerganov 2025-03-26 21:24:28 +02:00
  • 2447ad8a98 upgrade to llguidance 0.7.10 (#12576) b4964 Michał Moskal 2025-03-26 11:06:09 -07:00
  • 02082f1519 clip: Fix llama-llava-clip-quantize-cli quantization error under CUDA backend (#12566) b4963 Ivy233 2025-03-26 22:06:04 +08:00
  • df4d20cd53 convert : fix squeeze for ssm_conv tensors (#12573) Georgi Gerganov 2025-03-26 14:21:05 +02:00
  • 5ed38b6852 ggml : fix MUL_MAT_ID repack with Q8_K (#12544) b4961 Georgi Gerganov 2025-03-26 13:02:00 +02:00
  • fd7855f8f5 doc: [MUSA] minor changes (#12583) R0CKSTAR 2025-03-26 15:09:48 +08:00
  • 53af4dba42 convert: fix Mistral3/Gemma3 model hparams init (#12571) Sigbjørn Skjæret 2025-03-25 23:03:10 +01:00
  • c6a1be6c0b metal : fix typo [no ci] Georgi Gerganov 2025-03-25 22:37:56 +02:00
  • e7d14ab26c metal : reduce register pressure Georgi Gerganov 2025-03-25 21:55:15 +02:00
  • 20b256e0fd convert : match ssm_conv tensors by type gg/mamba-fix-squeeze Francis Couture-Harpin 2025-03-25 14:29:22 -04:00
  • 9c60fc4c78 convert : fix squeeze for ssm_conv tensors Georgi Gerganov 2025-03-25 19:54:18 +02:00
  • ef19c71769 run: de-duplicate fmt and format functions and optimize (#11596) b4958 Eric Curtin 2025-03-25 17:46:11 +00:00
  • fe12e20a7f metal : mv q6_K support nr0 > 1 Georgi Gerganov 2025-03-25 17:48:43 +02:00
  • 51dea76888 metal : fix nr constant [no ci] Georgi Gerganov 2025-03-25 14:40:01 +02:00
  • 982c82f1e6 metal : fix comments [no ci] Georgi Gerganov 2025-03-25 14:36:22 +02:00
  • 24a9ea8b44 metal : rename all_sum -> sum_all Georgi Gerganov 2025-03-25 14:34:38 +02:00
  • fcca45c027 metal : refactor mat-vec code Georgi Gerganov 2025-03-24 16:51:30 +02:00
  • 053b3f9aae ggml-cpu : update KleidiAI to v1.5.0 (#12568) b4957 Dan Johansson 2025-03-25 12:10:18 +01:00
  • e2f560175a SYCL: disable Q4_0 reorder optimization (#12560) b4956 Akarshan Biswas 2025-03-25 16:10:18 +05:30
  • 36ee06dd2d docs : add build instructions for KleidiAI (#12563) Dan Johansson 2025-03-25 10:35:20 +01:00
  • 3cd3a39532 ci: [MUSA] add CI and update doc (#12562) R0CKSTAR 2025-03-25 15:45:08 +08:00
  • e94c2bd360 ggml : improve repack templates gg/repack-fix-mul-mat-id Georgi Gerganov 2025-03-25 09:42:31 +02:00
  • 2d77d88e70 context : fix worst-case reserve outputs (#12545) b4953 Georgi Gerganov 2025-03-25 09:19:23 +02:00
  • b8b7885484 SYCL: disable Q4_0 reorder optimization sycl/disable_reorder_opt Akarshan Biswas 2025-03-25 10:14:11 +05:30
  • c95fa362b3 ci: [SYCL] ggml-ci Use main GPU and enable sysman (#12547) Akarshan Biswas 2025-03-24 23:05:38 +05:30
  • 2b65ae3029 opencl: simplify kernel embedding logic in cmakefile (#12503) b4951 lhez 2025-03-24 09:20:47 -07:00
  • 48d7021c61 CI: fix SYCL build (#12546) Akarshan Biswas 2025-03-24 18:28:32 +05:30
  • 87cd537a29 ggml : fix MUL_MAT_ID repack with Q8_K Georgi Gerganov 2025-03-24 13:07:10 +02:00
  • 3361e2deba docs: update: improve the Fedoa CUDA guide (#12536) Tei Home 2025-03-24 19:02:26 +08:00
  • 00d53800e0 llama-vocab : add SuperBPE pre-tokenizer (#12532) b4948 compilade 2025-03-24 06:47:24 -04:00
  • 7ea75035b6 CUDA: Fix clang warnings (#12540) b4947 R0CKSTAR 2025-03-24 18:28:34 +08:00
  • c54f6b7988 mmap : skip resource limit checks on AIX (#12541) b4946 Prajwal B Mehendarkar 2025-03-24 15:47:10 +05:30
  • 9b169a4d4e vulkan: fix mul_mat_vec failure in backend tests (#12529) b4945 Jeff Bolz 2025-03-24 01:56:17 -05:00
  • a5b1943912 ggml-quants : fix some edge cases in make_qkxh_nl_quants compilade/optimal-rounding Francis Couture-Harpin 2025-03-23 17:59:37 -04:00
  • 35c2f8b9ff llama-vocab : add SuperBPE pre-tokenizer compilade/superbpe Francis Couture-Harpin 2025-03-23 16:19:03 -04:00
  • 77f9c6bbe5 server : Add verbose output to OAI compatible chat endpoint. (#12246) b4944 Marius Gerdes 2025-03-23 19:30:26 +01:00
  • 18b663d8e4 install : add macports (#12518) Lars Sonchocky-Helldorf 2025-03-23 09:21:48 +01:00
  • 8b8b88f3de ggml-quants : restore Q2_K use of make_qp_quants Francis Couture-Harpin 2025-03-22 18:47:56 -04:00
  • fbdfefe74e llama : gemma3 : use output tensor if it exists in model weight (#12506) b4942 Xuan-Son Nguyen 2025-03-22 23:28:19 +01:00
  • a41139723d Merge branch 'master' into compilade/optimal-rounding Francis Couture-Harpin 2025-03-22 15:05:11 -04:00
  • af23abd3cb ggml-quants : remove slower qsort-based cumulative search Francis Couture-Harpin 2025-03-22 12:07:28 -04:00
  • 3e4b675c9f ggml-quants : use a max-heap for TQ1_0 and TQ2_0 quantization Francis Couture-Harpin 2025-03-22 12:03:26 -04:00
  • ba932dfb50 ggml : fix quantized cpy op (#12310) Georgi Gerganov 2025-03-22 16:23:26 +02:00
  • fac63a3d78 musa: refine compute capability (#12493) b4940 R0CKSTAR 2025-03-22 17:11:37 +08:00
  • eddfb43850 vulkan: Optimize mul_mat_vec p021 and nc shaders (#12505) b4939 Jeff Bolz 2025-03-22 03:40:11 -05:00
  • 4375415b4a Vulkan: RTE rounding for cpy to quant (#12480) b4938 stduhpf 2025-03-21 20:34:50 +01:00
  • 30c42ef5cb vulkan: workaround for AMD Windows driver 16 bit unpack8 bug (#12472) b4937 Eve 2025-03-21 19:27:47 +00:00
  • f86b8ff210 ggml-quants : use qkxh in more places Francis Couture-Harpin 2025-03-21 14:05:58 -04:00
  • af04481e6b model : do not repack if a GPU device is present (#12498) b4936 Georgi Gerganov 2025-03-21 16:14:29 +02:00
  • 960e726077 chore : cleanup llama_model_loader::TENSOR_ usage (#12492) b4935 Sigbjørn Skjæret 2025-03-21 10:21:36 +01:00
  • ea1518e839 llama-tts : avoid crashes related to bad model file paths (#12482) b4934 marcoStocchi 2025-03-21 10:12:45 +01:00
  • 1aa87ee53d [SYCL] Fix build on Windows when ccache enabled (#9954) (#9976) b4933 蕭澧邦 2025-03-21 14:58:47 +08:00
  • 9ffcc9e374 sycl: cleanup oneDNN related code (#12097) b4932 Svetlozar Georgiev 2025-03-21 02:15:56 +00:00
  • 3be115100f ggml-quants : use a max-heap for linear quants like Q3_K Francis Couture-Harpin 2025-03-20 19:21:45 -04:00
  • b8b173274d server : remove old commented code [no ci] xsn/private_batch_api_pooling_none Georgi Gerganov 2025-03-20 18:19:55 +02:00
  • e04643063b webui : Prevent rerendering on textarea input (#12299) Woof Dog 2025-03-20 14:57:43 +00:00
  • 8a23b4a54a server : avoid common_batch Georgi Gerganov 2025-03-20 16:52:24 +02:00
  • dbb3a4739e llama : make Qwen2MoE QKV bias optional (#12477) b4930 Sigbjørn Skjæret 2025-03-20 12:49:59 +01:00
  • 3d82dbcbce ggml : block interleaving support for Q4_K quantization for x86 AVX2 architecture (#12332) b4929 Srihari-mcw 2025-03-20 17:05:34 +05:30
  • 76fd7d6f5b perplexity : avoid common_batch Georgi Gerganov 2025-03-20 12:21:40 +02:00