Default Branch

f11d642a27 · HIP: avoid treating CDNA as dgx spark for gqa_ratio 20 in fattn_mma dqk 576 (#29572) · Updated 2026-10-01 00:51:38 -05:00

Branches

f9968f661d · ggml : update comments [no ci] · Updated 2024-09-11 05:16:39 -05:00    upstream-archive

7603
5

cfbf33a705 · ggml : style changes + fix 512-bit nb loop check · Updated 2024-09-09 04:50:35 -05:00    upstream-archive

7664
4

c3e2bb6dcf · rpc : fix nkvo · Updated 2024-09-06 20:24:47 -05:00    upstream-archive

7645
1

b979fc97ba · cmake : use ggml-metal.metal from source dir to build default.metallib · Updated 2024-09-05 11:17:56 -05:00    upstream-archive

7654
1

75b3a09602 · test-backend-ops : add TQ1_0 and TQ2_0 comments for later · Updated 2024-09-04 14:00:21 -05:00    upstream-archive

7656
33

f648ca2cee · llama : add llama_sampling API + move grammar in libllama · Updated 2024-09-03 02:31:54 -05:00    upstream-archive

7663
1

40fa68cb46 · readme : add API change notice · Updated 2024-09-02 10:32:24 -05:00    upstream-archive

7672
3

a95225cdfd · metal : another fix for the fa kernel · Updated 2024-08-26 07:08:38 -05:00    upstream-archive

7696
1

aa931d0375 · metal : fix fa kernel · Updated 2024-08-26 05:09:50 -05:00    upstream-archive

7696
1

6494509801 · backup · Updated 2024-08-26 03:58:54 -05:00    upstream-archive

7706
2

ccb45186d0 · docs : remove references · Updated 2024-08-26 01:52:17 -05:00    upstream-archive

7700
2

8062650343 · llama : fix simple splits when the batch contains embeddings · Updated 2024-08-21 14:09:03 -05:00    upstream-archive

7711
19

9127800d83 · wip · Updated 2024-08-16 18:51:06 -05:00    upstream-archive

7744
2

62d7b6c87f · cuda : re-add q4_0 · Updated 2024-08-14 05:37:03 -05:00    upstream-archive

7740
3

93ec58b932 · server : fix typo in comment · Updated 2024-08-13 21:12:26 -05:00    upstream-archive

7742
4

faaac59d16 · llama : support NUL bytes in tokens · Updated 2024-08-11 20:00:03 -05:00    upstream-archive

7753
1

73bc9350cd · gguf-py : Numpy dequantization for grid-based i-quants · Updated 2024-08-09 22:47:31 -05:00    upstream-archive

7773
2

9329953a61 · llama : avoid double tensor copy when saving session to buffer · Updated 2024-08-07 15:03:34 -05:00    upstream-archive

7781
2

cad8abb49b · add tool to allow plotting tensor allocation maps within buffers · Updated 2024-08-06 15:09:51 -05:00    upstream-archive

7747
1

6e299132e7 · clip : style changes · Updated 2024-08-06 03:44:29 -05:00    upstream-archive

8071
56