Default Branch

a7a6d0d269 · vulkan: extend topk_moe fusion to support sqrt(softplus) (#26124) · Updated 2026-08-01 14:18:07 -05:00

Branches

9998ecd191 · llama : add phixtral support (wip) · Updated 2024-01-13 06:24:07 -06:00    upstream-archive

8373
1

1fb563ebdc · py : try to fix flake stuff · Updated 2024-01-13 05:42:35 -06:00    upstream-archive

8374
2

9bfcb16fd3 · Add llama enum for IQ2_XS · Updated 2024-01-11 10:24:12 -06:00    upstream-archive

8423
11

24096933b0 · server : try to fix infill when prompt is empty · Updated 2024-01-09 03:27:29 -06:00    upstream-archive

8425
1

7216af5c09 · ggml : fix 32-bit ARM compat (cont) · Updated 2024-01-09 02:33:16 -06:00    upstream-archive

8428
2

d57cb9c294 · passkey : add readme · Updated 2024-01-08 03:13:44 -06:00    upstream-archive

8438
7

7cfde78190 · llama : remove redundant GQA check · Updated 2024-01-06 08:04:20 -06:00    upstream-archive

8446
1

9f51f3e695 · metal : opt mul_mm_id · Updated 2024-01-02 12:50:18 -06:00    upstream-archive

8472
17

4cc78d3873 · ggml : force F32 precision for ggml_mul_mat · Updated 2024-01-02 09:54:56 -06:00    upstream-archive

8471
1

b5af7ad84f · llama : refactor quantization to avoid <mutex> header · Updated 2024-01-02 07:56:57 -06:00    upstream-archive

8474
1

120a1a5515 · llama : auto download HF models if URL provided · Updated 2024-01-02 05:29:06 -06:00    upstream-archive

8475
1

f64e4f04e7 · ggml : testing GPU FP precision via quantized CPY · Updated 2023-12-30 11:11:40 -06:00    upstream-archive

8493
1

f32f30bc57 · test · Updated 2023-12-26 09:52:42 -06:00    upstream-archive

8523
1

ab1b75166f · Merge branch 'master' into gg/ggml_scale · Updated 2023-12-21 14:35:11 -06:00    upstream-archive

8546
4

7c87353e61 · common : remove incorrect --model-draft default · Updated 2023-12-21 11:17:12 -06:00    upstream-archive

8554
1

a40f6110f0 · ggml : force F32 precision for ggml_mul_mat · Updated 2023-12-19 08:34:59 -06:00    upstream-archive

8561
1

3c734f4941 · plamo : testing · Updated 2023-12-18 09:06:05 -06:00    upstream-archive

8566
13

a462159c43 · cuda : ggml_cuda_op_mul_mat_cublas support F32 precision · Updated 2023-12-18 06:24:29 -06:00    upstream-archive

8566
16

1b05817112 · decode : fix logits_valid for old API · Updated 2023-12-17 17:49:21 -06:00    upstream-archive

8567
1

865066621b · llama.swiftui : improve bench · Updated 2023-12-17 11:37:22 -06:00    upstream-archive

8581
12