Commit Graph

  • 54ef9cfc72 vulkan: Throttle the number of shader compiles during the build step. (#10222) b4067 Jeff Bolz 2024-11-11 11:13:51 -06:00
  • b0cefea58a metal : more precise Q*K in FA vec kernel (#10247) b4066 Georgi Gerganov 2024-11-11 08:39:13 +02:00
  • b141e5f6ef server : enable KV cache defrag by default (#10233) b4065 Georgi Gerganov 2024-11-11 08:38:43 +02:00
  • 4b3a9212b6 flake.lock: Update (#10243) Georgi Gerganov 2024-11-10 21:45:25 +02:00
  • 505f33274d server : (web UI) Add back sampler settings (#10239) MaggotHATE 2024-11-11 00:42:25 +05:00
  • 160687b3ed vulkan: Fix newly added tests for permuted mul_mat and 1D im2col (#10226) b4062 Jeff Bolz 2024-11-10 05:37:56 -06:00
  • 6423c65aa8 metal : reorder write loop in mul mat kernel + style (#10231) b4061 Georgi Gerganov 2024-11-09 11:53:13 +02:00
  • 39a334a9aa metal : fix build and some more comments (#10229) b4060 Georgi Gerganov 2024-11-09 11:53:02 +02:00
  • bb38cdd8ba metal : fix F32 accumulation in FA vec kernel (#10232) b4059 Georgi Gerganov 2024-11-09 11:52:45 +02:00
  • f018acba22 llama : fix Qwen model type strings b4058 Georgi Gerganov 2024-11-09 11:26:34 +02:00
  • 46323fa9ef metal : hide debug messages from normal log b4057 Georgi Gerganov 2024-11-09 11:21:49 +02:00
  • 3d1fe1bb4d metal : int -> short, style gg/metal-mul-mat-write-opt Georgi Gerganov 2024-11-08 15:38:25 +02:00
  • 535050572a metal : reorder write loop Georgi Gerganov 2024-11-08 15:15:25 +02:00
  • bd1198a67a metal : fix build and some more comments gg/metal-fix-build Georgi Gerganov 2024-11-09 10:09:50 +02:00
  • 5b359bb1e3 ggml: fix zero division in ‘dne’ calculation in CUDA COUNT_EQUAL operator when ‘ne’ is small (#10213) b4056 SXX 2024-11-09 15:35:46 +08:00
  • e89213492d ggml : optimize llamafile cpu matrix multiplication for ppc64le (#10156) b4055 amritahs-ibm 2024-11-09 12:47:50 +05:30
  • 8fc393f246 scripts : fix pattern and get n_tokens in one go (#10221) haopeng 2024-11-09 15:06:54 +08:00
  • ec450d3bbf metal : opt-in compile flag for BF16 (#10218) b4053 Georgi Gerganov 2024-11-08 21:59:46 +02:00
  • 695ad752b2 metal : improve clarity (minor) (#10171) b4052 Georgi Gerganov 2024-11-08 18:37:41 +02:00
  • 841f27abdb metal : optimize FA kernels (#10171) Georgi Gerganov 2024-11-08 13:47:22 +02:00
  • a2385da59c make : clean-up [no ci] gg/metal-fa-f16 Georgi Gerganov 2024-11-08 13:46:20 +02:00
  • d05b3127bd swift : exclude ggml-metal-embed.metal (#10211) b4050 Jhen-Jie Hong 2024-11-08 17:34:06 +08:00
  • b89e71b195 metal : fix BF16 requirement for FA kernels Georgi Gerganov 2024-11-08 11:28:04 +02:00
  • bc143ecf81 cuda : disable BF16 FA Georgi Gerganov 2024-11-08 10:27:43 +02:00
  • 5d1a10d275 metal : prevent int overflows [no ci] Georgi Gerganov 2024-11-07 22:11:24 +02:00
  • 486a5eb8c1 build : remove obsolete compile flag [no ci] Georgi Gerganov 2024-11-07 21:51:28 +02:00
  • 120d51285c metal : compile-guard bf16 FA kernels Georgi Gerganov 2024-11-07 21:38:37 +02:00
  • 2fccc8ac2d metal : minor clean-up Georgi Gerganov 2024-11-07 21:29:22 +02:00
  • 7facc29d69 metal : use F16 precision in FA kernels Georgi Gerganov 2024-11-06 15:33:30 +02:00
  • 25e877309a ggml : add ggml_flash_attn_ext_get_prec Georgi Gerganov 2024-11-06 15:09:47 +02:00
  • 76c6e7f105 server : minor UI fix (#10207) Xuan Son Nguyen 2024-11-07 18:44:38 -04:00
  • a71d81cf8c server : revamp chat UI with vuejs and daisyui (#10175) b4048 Xuan Son Nguyen 2024-11-07 17:31:10 -04:00
  • eec4d71737 scripts : add amx to sync-ggml.sh [no ci] Georgi Gerganov 2024-11-07 23:11:36 +02:00
  • 3b08828674 sync : ggml Georgi Gerganov 2024-11-07 23:08:24 +02:00
  • a2c6fd747c scripts : sync update Georgi Gerganov 2024-11-07 23:07:55 +02:00
  • 94accca4c2 vec move mask to shmem gg/metal-fa-f16-save Georgi Gerganov 2024-11-07 20:58:10 +02:00
  • 3b9625032c f16 vec Georgi Gerganov 2024-11-07 20:34:16 +02:00
  • 8f0ef15265 clean-up Georgi Gerganov 2024-11-07 20:02:31 +02:00
  • 022e5e90e9 remove compile flag Georgi Gerganov 2024-11-07 19:18:31 +02:00
  • 97404c4a03 ggml : add ggml-cpu.h to the public headers (#10204) b4044 Diego Devesa 2024-11-07 18:16:08 +01:00
  • 60e17ce23c Remove identical wte/etw logic for jais (#10203) Faisal Zaghloul 2024-11-07 11:46:12 -05:00
  • a6c8dbfa5d wip Georgi Gerganov 2024-11-07 18:20:25 +02:00
  • 5107e8cea3 DRY: Fixes clone functionality (#10192) b4042 wwoodsTM 2024-11-07 08:20:25 -07:00
  • 4abeb60a1a int64 dst Georgi Gerganov 2024-11-07 17:17:29 +02:00
  • 3ab47eb746 float -> half regs Georgi Gerganov 2024-11-07 17:06:34 +02:00
  • e121d82f6a 64-bit -> 32-bit Georgi Gerganov 2024-11-07 17:00:06 +02:00
  • a75cdcca60 remove inner if mask Georgi Gerganov 2024-11-07 16:40:29 +02:00
  • 61d05b57d9 remove ms array Georgi Gerganov 2024-11-07 13:35:33 +02:00
  • 984928109c move mask to shared mem Georgi Gerganov 2024-11-07 13:12:10 +02:00
  • 2aedbb354e wip 5 Georgi Gerganov 2024-11-07 12:32:59 +02:00
  • dc2a27f2a2 wip 4 Georgi Gerganov 2024-11-07 09:26:10 +02:00
  • 2319126a70 fix q4_0_8_8 format for corrupted tokens issue (#10198) b4041 snadampal 2024-11-07 02:02:08 -06:00
  • 3bcd40b3c5 Optimize RWKV6 Operator Naming and Implement Multi-core CPU/ SYCL Acceleration (#10133) b4040 Zhiyuan Li 2024-11-07 18:19:10 +11:00
  • 9bd5ae09ae wip 3 Georgi Gerganov 2024-11-06 22:52:33 +02:00
  • 2335086fd3 wip2 Georgi Gerganov 2024-11-06 22:04:07 +02:00
  • 01c7f11224 wip Georgi Gerganov 2024-11-06 21:06:56 +02:00
  • 0f7e8f389d metal : add GGML_METAL_FORCE_FATTN_PREC_F16 Georgi Gerganov 2024-11-06 16:21:37 +02:00
  • eefc132bb7 metal : use F16 precision in FA kernel Georgi Gerganov 2024-11-06 15:33:30 +02:00
  • 22a9311a1a ggml : add ggml_flash_attn_ext_get_prec Georgi Gerganov 2024-11-06 15:09:47 +02:00
  • 5c333e0140 metal : add BF16 support (#8439) Georgi Gerganov 2024-11-06 19:53:51 +02:00
  • b11f9ba9b8 server : remove hack for extra parallel slot (#10187) b4038 Georgi Gerganov 2024-11-06 13:29:01 +02:00
  • 94d8cb8be1 metal : fix from ptr buffer name (#10189) b4037 Diego Devesa 2024-11-06 12:10:07 +01:00
  • 1dc04b2dee ggml : adjust is_first_call init value (#10193) b4036 Georgi Gerganov 2024-11-06 11:20:10 +02:00
  • a1eaf6a960 metal : add quantized FA support (#10149) Georgi Gerganov 2024-11-06 10:24:23 +02:00
  • c5d8bb5a81 leave only basic functions for SYCL CI fix_sycl_ci Meng, Hengyu 2024-11-06 07:47:50 +00:00
  • b8deef0ec0 llama : add <|tool_call|> formatting to Granite template (#10177) b4034 Gabe Goodhart 2024-11-05 05:23:04 -07:00
  • a9e8a9a030 ggml : fix arch check in bf16_to_fp32 (#10164) b4033 Diego Devesa 2024-11-04 23:17:01 +01:00
  • 3407364776 Q6_K AVX improvements (#10118) b4032 Eve 2024-11-04 22:06:31 +00:00
  • b4e9c5998d convert : fix flake8 lint Francis Couture-Harpin 2024-11-04 15:26:15 -05:00
  • 8d8f065743 Merge branch 'master' into compilade/mamba2 Francis Couture-Harpin 2024-11-04 14:30:18 -05:00
  • d5a409e57f ggml : fix gelu tables initialization (#10172) Diego Devesa 2024-11-04 20:06:58 +01:00
  • 3bc7103d2e ggml : avoid multiply by D in GGML_OP_SSM_SCAN Francis Couture-Harpin 2024-11-04 11:36:37 -05:00
  • 401558b7ba ggml : fix q4xx mat mul, increase ggml_aligned_malloc alignment (#10167) Diego Devesa 2024-11-04 17:34:08 +01:00
  • 9e0ecfb697 server : clarify /slots endpoint, add is_processing (#10162) Xuan Son Nguyen 2024-11-04 16:33:29 +01:00
  • 6a066b9978 fix build break on arm64 linux (#10166) snadampal 2024-11-04 09:08:33 -06:00
  • ea02c753eb cuda : clear error after changing peer access (#10153) b4027 Diego Devesa 2024-11-04 13:10:23 +01:00
  • 05697f670b metal : simplify f16 and f32 dequant kernels (#0) b4026 Georgi Gerganov 2024-11-04 13:49:34 +02:00
  • f8e58135cf metal : move dequantize templates to beginning of MSL source (#0) b4025 Georgi Gerganov 2024-11-04 13:43:32 +02:00
  • 329ed914c9 CANN: adjust backend registry refactor. (#10158) b4024 leo-pony 2024-11-04 19:08:22 +08:00
  • ce027adfb3 sync : ggml b4023 Georgi Gerganov 2024-11-04 10:33:37 +02:00
  • 284e5b0275 cmake : make it possible linking ggml as external lib (ggml/1003) Yuri Khrustalev 2024-11-02 05:09:12 -04:00
  • e2292aaa17 metal : fix minor string leaks (ggml/1004) Plamen Minev 2024-11-01 16:55:10 +02:00
  • 9f40989351 ggml : move CPU backend to a separate file (#10144) b4020 Diego Devesa 2024-11-03 19:34:08 +01:00
  • 08828a6d7d metal : minor fixup in FA kernel (#10143) b4019 Georgi Gerganov 2024-11-03 15:18:40 +02:00
  • 1839f69130 flake.lock: Update (#10146) Georgi Gerganov 2024-11-03 15:14:15 +02:00
  • 9830b6923b Add apple arm to presets (#10134) Christian Köhnenkamp 2024-11-02 23:35:31 +01:00
  • 42cadc74bd server : fix slot selection by lru (#10126) b4016 sasha0552 2024-11-02 16:34:56 +00:00
  • 45950415ed server : fix endpoint checks (#10135) b4015 Georgi Gerganov 2024-11-02 18:34:00 +02:00
  • 4fc8673d09 llama-bench : skip repeated values in consecutive lines sl/llama-bench-headers slaren 2024-11-02 15:37:33 +01:00
  • 1926d6e39d llama : adjust default context size + print warnings (#10136) b4014 Georgi Gerganov 2024-11-02 15:18:56 +02:00
  • b634f8a26f simple-chat : only add bos on first prompt (#10129) b4013 Diego Devesa 2024-11-02 13:08:53 +01:00
  • 7554aa4655 convert-lora : make --base optional (#10110) Xuan Son Nguyen 2024-11-02 12:53:17 +01:00
  • 20e12112fd llama : suggest reduce ctx size when kv init fails sl/aligned-alloc-no-abort slaren 2024-11-02 00:55:19 +01:00
  • bf60f27cda ggml : do not abort when ggml_aligned_malloc fails slaren 2024-11-02 00:54:16 +01:00
  • a6744e43e8 llama : add simple-chat example (#10124) b4011 Diego Devesa 2024-11-01 23:50:59 +01:00
  • e991e3127f llama : use smart pointers for ggml resources (#10117) b4010 Diego Devesa 2024-11-01 23:48:26 +01:00
  • 418f5eef26 vulkan : improve ggml_vk_create_buffer error handling (#9898) b4009 Shupei Fan 2024-11-02 02:33:14 +08:00
  • ba6f62eb79 readme : update hot topics Georgi Gerganov 2024-11-01 17:31:51 +02:00
  • 7d16e1bc8c Merge branch 'master' into compilade/mamba2 Francis Couture-Harpin 2024-11-01 11:12:18 -04:00
  • d865d1478c server : fix smart selection of available slot (#10120) b4007 sasha0552 2024-11-01 13:33:14 +00:00