mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-28 17:07:31 -05:00
b11224
* vulkan: read the batch stride of an in place src0 from nb[2] A dim01 contiguous tensor can still be a view whose batches are strided by more than ne[1] rows, the first rows of a KV cache for example. Both the mat-vec and the matrix paths read such a tensor in place but passed ne00*ne01 as the batch stride, so every head past the first read the wrong rows. The same applies to src1. The stride now comes from nb[2] whenever the tensor is used in place; the value is unchanged for a contiguous tensor. test-backend-ops gets an m_v parameter on test_mul_mat, the number of rows of a in memory, and two cases at the shapes of a decoder self attention over a cache. * vulkan: size the in place A and B ranges by their strided extent The matrix path bound src0 and src1 to the shader with a range of elements times type size, which ends before the batches of a strided view. Pipelines with bounded access read zero past that range, so the same view that the mat-vec path already handles gave wrong results on Intel and on NVIDIA without coopmat2. The range now comes from ggml_nbytes when the tensor is read in place. * vulkan: address review from jeffbolznv Bind the in place A and B of the matrix path with ggml_vk_subbuffer, which spans to the end of the buffer, so a strided view is in range without computing its extent. mul_mat_id reads the batch stride of an in place src0 and src1 with the same helper as mul_mat. test_mul_mat_id gets an m_v parameter, the number of rows of as in memory, and a case whose experts are strided by more rows than it uses. * vulkan: read the batch stride of an in place src0 in mul_mat_vec_id The single token path of mul_mat_id passed ne00*ne01 as the batch stride of A, so a strided expert view read the wrong rows. The stride now comes from ggml_vk_batch_stride like the other three paths, and src1 follows the same rule. test_mul_mat_id gets a single token case over the strided view. * vulkan: address review from jeffbolznv The batch stride of an in place tensor is taken from nb[2] as nb[2] / type_size * block_size, which holds when nb[2] is padded and not a multiple of nb[1]. A test_mul_mat case with a padded batch stride covers it. * vulkan: keep the A and B ranges exact in mul_mm The quantized A loads of mul_mm carry no row bound and rely on the descriptor range to read zeros past the last row of a partial tile. Binding A and B up to the end of the buffer let those tiles read the leftovers of a previous node and hung the NVFP4 mul_mm on NVIDIA without coopmat2. The range is the strided extent of a tensor read in place and the staged size otherwise.
server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)
server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)
server : allow RANK pooling batch splitting for causal LLM rerankers (ie. Qwen3 and Qwen3-VL) (#28876)
llama.cpp
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain
Languages
C++
55.7%
C
16.2%
Python
7.3%
Cuda
5.4%
TypeScript
4.1%
Other
11.1%