Default Branch

a7a6d0d269 · vulkan: extend topk_moe fusion to support sqrt(softplus) (#26124) · Updated 2026-08-01 14:18:07 -05:00

Branches

49a483e0f2 · wip · Updated 2024-02-04 04:34:36 -06:00    upstream-archive

8182
60

a647257b47 · cuda : express strides with helper constants · Updated 2024-02-04 03:45:26 -06:00    upstream-archive

8182
60

b957b8f5f6 · cuda : add flash_attn kernel (wip) · Updated 2024-02-01 11:49:57 -06:00    upstream-archive

8186
39

ac26f27028 · cuda : increase C to 128 for better performance · Updated 2024-02-01 09:08:29 -06:00    upstream-archive

8186
61

1ad42b1f1e · ggml : ggml_soft_max uses F16 mask · Updated 2024-01-31 12:33:59 -06:00    upstream-archive

8186
36

719a087138 · iq3_xxs: forgotten update of the grid points · Updated 2024-01-30 10:39:07 -06:00    upstream-archive

8200
1

2bf91c5306 · metal : clean up · Updated 2024-01-25 05:29:45 -06:00    upstream-archive

8296
23

6ccbd1777a · wip · Updated 2024-01-24 07:45:04 -06:00    upstream-archive

8296
18

da23b56f25 · wip : no ic 8 step · Updated 2024-01-24 05:25:34 -06:00    upstream-archive

8296
18

06c2d0d117 · wip · Updated 2024-01-23 14:42:43 -06:00    upstream-archive

8296
14

a9681febd6 · ggml : online attention (CPU) · Updated 2024-01-20 08:45:41 -06:00    upstream-archive

8296
4

32a392fe68 · try a differerent fix · Updated 2024-01-19 16:10:23 -06:00    upstream-archive

8297
2

4a3bc1522e · py : linting with mypy and isort · Updated 2024-01-19 14:18:58 -06:00    upstream-archive

8298
3

1453215165 · kompute : fix ggml_add kernel · Updated 2024-01-18 16:09:16 -06:00    upstream-archive

8414
105

ccc78a200e · hellaswag: speed up even more by parallelizing log-prob evaluation · Updated 2024-01-18 10:25:29 -06:00    upstream-archive

8314
1

2917e6b528 · Merge branch 'master' into gg/imatrix-gpu-4931 · Updated 2024-01-17 10:43:45 -06:00    upstream-archive

8321
10

23742deb5b · py : fix padded dummy tokens (I hope) · Updated 2024-01-17 07:44:22 -06:00    upstream-archive

8340
4

9fd1e83f6d · Use Q4_K for attn_v for Q2_K_S when n_gqa >= 4 · Updated 2024-01-17 04:16:08 -06:00    upstream-archive

8326
1

49bafe0986 · tests : avoid creating RNGs for each tensor · Updated 2024-01-17 02:40:55 -06:00    upstream-archive

8329
6

bb9abb5cd8 · imatrix: guard Q4_0/Q5_0 against ffn_down craziness · Updated 2024-01-16 01:56:05 -06:00    upstream-archive

8343
2