Default Branch

a7a6d0d269 · vulkan: extend topk_moe fusion to support sqrt(softplus) (#26124) · Updated 2026-08-01 14:18:07 -05:00

Branches

f86b9d152c · lookup : minor · Updated 2023-12-17 09:25:28 -06:00    upstream-archive

8579
9

d2f1e0dacc · Merge branch 'cuda-cublas-opts' into gg/phi-2 · Updated 2023-12-17 00:41:46 -06:00    upstream-archive

8577
17

b0547d2196 · gguf-py : fail fast on nonsensical special token IDs · Updated 2023-12-15 17:06:42 -06:00    upstream-archive

8579
1

c8554b80be · Merge branch 'master' of https://github.com/ggerganov/llama.cpp into ceb/fix-cuda-warning-flags · Updated 2023-12-13 11:06:01 -06:00    upstream-archive

8591
12

e1241d9b46 · metal : switch to execution barriers + fix one of the barriers · Updated 2023-12-13 05:56:45 -06:00    upstream-archive

8602
47

fc5f334689 · readme : add API change notice · Updated 2023-12-07 04:35:02 -06:00    upstream-archive

8604
15

af99c6fbfc · llama : remove memory_f16 and kv_f16 flags · Updated 2023-12-05 10:18:16 -06:00    upstream-archive

8616
26

3cb1c348b3 · metal : try to improve batched decoding · Updated 2023-12-01 14:01:58 -06:00    upstream-archive

8621
2

eb594c0f7d · alloc : fix build with debug · Updated 2023-12-01 02:46:05 -06:00    upstream-archive

8645
14

5b74310e6e · build : enable libstdc++ assertions for debug builds · Updated 2023-11-30 17:18:24 -06:00    upstream-archive

8630
1

bb39b87964 · ggml : restore abort() in GGML_ASSERT · Updated 2023-11-27 18:27:09 -06:00    upstream-archive

8649
1

87f4102a70 · llama : revert n_threads_batch logic · Updated 2023-11-27 13:47:35 -06:00    upstream-archive

8650
3

6272b6764a · use stride=128 if built for tensor cores · Updated 2023-11-27 12:09:14 -06:00    upstream-archive

8653
3

8d8b76d469 · lookahead : add comments · Updated 2023-11-26 03:26:55 -06:00    upstream-archive

8665
9

21b70babf7 · straightforward /v1/models endpoint · Updated 2023-11-24 10:22:39 -06:00    upstream-archive

8666
12

f8e9f11428 · common : add -dkvc arg for enabling kv cache dumps · Updated 2023-11-23 10:47:56 -06:00    upstream-archive

8672
4

f824902623 · YaRN : correction to GPT-NeoX implementation · Updated 2023-11-15 16:10:52 -06:00    upstream-archive

8704
1

d0445a2eff · better documentation · Updated 2023-11-09 18:38:20 -06:00    upstream-archive

8721
3

47d604fa2d · fix issues · Updated 2023-11-05 06:20:22 -06:00    upstream-archive

8735
3

3ef358fffd · Revert "cuda : use CUDA memory pool with async memory allocation/deallocation when available (#3903)" · Updated 2023-11-04 15:26:51 -05:00    upstream-archive

8739
2