Default Branch

d775ebf363 · server: return HTTP 400 for invalid embedding requests (#29060) · Updated 2026-10-01 09:52:43 -05:00

Branches

86ccd30983 · ci : only show warnings and errors in python type-check · Updated 2024-07-07 13:10:42 -05:00    upstream-archive

8002
10

a44f22e7d3 · py : use cpu-only torch in requirements.txt · Updated 2024-07-06 10:18:03 -05:00    upstream-archive

8012
1

f55b647300 · llama : minor indentation during tensor loading · Updated 2024-07-04 11:34:04 -05:00    upstream-archive

8034
16

dcab343f2f · use 1 seq for kl_divergence · Updated 2024-07-03 09:22:58 -05:00    upstream-archive

8049
2

703764a382 · convert : use non-fast T5 tokenizer · Updated 2024-07-02 11:29:26 -05:00    upstream-archive

8091
10

d4a1923d4e · minor : remove parentheses · Updated 2024-07-01 06:45:55 -05:00    upstream-archive

8069
2

51f0bd50a1 · Remove custom pre attention scaling and use computed value instead. · Updated 2024-06-29 22:02:50 -05:00    upstream-archive

8072
10

712e4d9450 · Generate full token count during warm up · Updated 2024-06-28 07:29:00 -05:00    upstream-archive

8075
1

65f9293d14 · devops : remove clblast + LLAMA_CUDA -> GGML_CUDA · Updated 2024-06-26 11:17:26 -05:00    upstream-archive

8101
1

1e6e363d7f · test zero max buffer size · Updated 2024-06-26 10:11:09 -05:00    upstream-archive

8102
1

ff0aa3abd1 · fix part of mul_mat_id · Updated 2024-06-20 22:38:00 -05:00    upstream-archive

8145
1

f3974cabac · all matrix multiplication backend · Updated 2024-06-18 06:18:26 -05:00    upstream-archive

8186
1

ce6e28cc23 · Update ggml-sycl.cpp · Updated 2024-06-18 03:57:14 -05:00    upstream-archive

8197
6

ef79941ac9 · llama : disable FA if KV head size do not match · Updated 2024-06-17 11:20:24 -05:00    upstream-archive

8167
1

a235b7c532 · Vectorize q load · Updated 2024-06-17 04:30:40 -05:00    upstream-archive

8197
11

98f948b9d0 · unicode : avoid char32_t · Updated 2024-06-16 05:18:46 -05:00    upstream-archive

8182
1

28f7a4d028 · ggml : fix handling of zero blocks in IQ quants · Updated 2024-06-16 02:41:53 -05:00    upstream-archive

8183
1

e9f2abfc8c · bitnet : pad tensors to 256 · Updated 2024-06-15 11:01:03 -05:00    upstream-archive

8201
25

34bdbed481 · rpc : fix load/store misaligned addresses · Updated 2024-06-15 06:39:20 -05:00    upstream-archive

8185
1

eaf34ba0cd · metal : utilize max shared memory for mul_mat_id · Updated 2024-06-14 05:02:25 -05:00    upstream-archive

8192
1