Default Branch

c85b92c69c · tests : adjust server string regex to also match m2 utlra results (#29648) · Updated 2026-09-29 07:32:35 -05:00

Branches

f824902623 · YaRN : correction to GPT-NeoX implementation · Updated 2023-11-15 16:10:52 -06:00    upstream-archive

9723
1

d0445a2eff · better documentation · Updated 2023-11-09 18:38:20 -06:00    upstream-archive

9740
3

47d604fa2d · fix issues · Updated 2023-11-05 06:20:22 -06:00    upstream-archive

9754
3

3ef358fffd · Revert "cuda : use CUDA memory pool with async memory allocation/deallocation when available (#3903)" · Updated 2023-11-04 15:26:51 -05:00    upstream-archive

9758
2

46868a499e · metal : multi-simd softmax · Updated 2023-11-01 14:16:34 -05:00    upstream-archive

9783
1

a8796f9609 · llm : cleanup + comments · Updated 2023-11-01 13:08:02 -05:00    upstream-archive

9792
4

7420bef83e · wip wip wip · Updated 2023-11-01 01:51:43 -05:00    upstream-archive

9792
1

afb3929279 · Merge branch 'master' into llama-refactor · Updated 2023-10-31 13:35:31 -05:00    upstream-archive

9794
21

29fe516913 · wip · Updated 2023-10-31 11:36:37 -05:00    upstream-archive

9795
1

dab42893c9 · scripts : working curl pipe · Updated 2023-10-31 10:03:56 -05:00    upstream-archive

9795
3

7923b70cb8 · llama : add llm_build_inp_embd helper · Updated 2023-10-31 09:43:08 -05:00    upstream-archive

9800
37

4b3cb98d46 · ggml-impl : move extern "C" to start of file · Updated 2023-10-30 12:05:58 -05:00    upstream-archive

9789
7
lto

bc28aaa8c2 · make : use -lfto=auto to avoid warnings and maintain perf · Updated 2023-10-30 09:00:53 -05:00    upstream-archive

9789
5

15267192c0 · llama : refactor tensor offloading as callback · Updated 2023-10-29 06:04:36 -05:00    upstream-archive

9793
15

8a86b95e87 · quantize : --pure option for disabling k-quant mixtures · Updated 2023-10-28 15:37:03 -05:00    upstream-archive

9794
3

de7e0912b6 · convert : ignore tokens if their IDs are within [0, vocab_size) · Updated 2023-10-28 07:01:36 -05:00    upstream-archive

9797
1

bbfc62ac2f · sampling : temp == 0.0 -> no probs, temp < 0.0 -> probs · Updated 2023-10-28 06:04:57 -05:00    upstream-archive

9805
3

cd3e20fb50 · cuda : fix multi-gpu with tensor cores · Updated 2023-10-27 15:11:50 -05:00    upstream-archive

9804
3

49af767fad · build : add compile option to force use of MMQ kernels · Updated 2023-10-27 05:21:04 -05:00    upstream-archive

9806
7

d798a17c34 · cuda : add TODO for calling cublas from kernel + using mem pool · Updated 2023-10-24 08:33:24 -05:00    upstream-archive

9820
10