mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-29 09:27:33 -05:00
b11243
* ci : update the oneAPI toolkit to 2026.1 oneDNN is removed from Intel Deep Learning Essentials in 2026.0, so staying on the deep-learning-essentials path would silently lose oneDNN support when the toolkit version is updated. Switch both the Ubuntu and Windows CI jobs to the new unified Intel oneAPI Toolkit installer, which still includes oneDNN (until 2027.0) and keeps the component IDs unchanged for the Windows install script. Measured with the same code (b10899) built with oneAPI 2026.1 vs the 2025.3-based release build on Arc B570: prompt processing 1331 vs 434 t/s (3.1x), token generation 50.1 vs 45.3-48.0 t/s. Assisted-by: GLM (z-ai/glm-5.3-flash) * docs : update the SYCL backend build requirements for oneAPI 2026.1 With the 2026.0 release the Base toolkit and the HPC toolkit are combined into the oneAPI Toolkit, and oneDNN is removed from the Deep Learning Essentials package. Update the install instructions, the verified release table and the news section accordingly. Assisted-by: GLM (z-ai/glm-5.3-flash) * ci : update the release workflow for oneAPI 2026.1 and Level Zero SDK 1.33.1 Align the release package build with the CI build update: - oneAPI toolkit 2025.3.3 -> 2026.1 (the unified oneAPI Toolkit) - Level Zero SDK 1.28.2 -> 1.33.1, and the Debian package names (level-zero/level-zero-devel -> libze1/libze-dev) - The Windows DLL copy list for the 2026.1 runtime: sycl9.dll and the .6/.3 MKL library versions Assisted-by: GLM (z-ai/glm-5.3-flash) * ci : remove the removed .spv fallback files from the Windows DLL copy list oneAPI 2026.1 no longer ships libsycl-fallback-bfloat16.spv and libsycl-native-bfloat16.spv (the OpenCL fallback mechanism changed), so the copy step failed with exit 1. Assisted-by: GLM (z-ai/glm-5.3-flash) * devops : update the oneAPI toolkit image in the Intel Dockerfile Assisted-by: GLM (z-ai/glm-5.3-flash) --------- Co-authored-by: Asahi-Prv <Asahi-Prv@users.noreply.github.com>
…
llama.cpp
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain
Languages
C++
55.7%
C
16.2%
Python
7.3%
Cuda
5.4%
TypeScript
4.1%
Other
11.1%