mirror of
https://github.com/ggml-org/llama.cpp.git
synced 2026-09-28 00:47:25 -05:00
* Added tiled mul_mat.
For each mul_mat_one_chunk, quants are unpacked into (max) 256x256 tiles of int8,
one routine per quent. Then microkernel computes 16x16 tiles before writing out
256x256 float reults to main memory.
Tests/benches in tests/test-tiled-mulmat.cpp. 3-6x speed improvement
for large matmul, break even at 4096x64 * 64x4096, 80% performance (net
loss) for GEMV. Error rates trivial (order of 1-e04 max, 1-e05 rmse).
* Fixes for ARM/windows builds
* more windows fixes, ggml-cpu.h isn't visible in MSVC for some reason
* unified iqp + tiled on the Q5_K, IQ4_XS set for benchmarking, updated benchmark
* Fixed accidental removal of llama_build_and_test(test-backend-ops.cpp)
* First integration of iqp code
Co-authored-by Bartowski <3266127+bartowski1182@users.noreply.github.com>
* Cleaning up declaration of iq unpacking helpers to align with the bit unpackers
* Removed iqp path
* Fix cross-platform warnings
* Disabling benchmarks unless explicitly enabled
* Fix backend_init for DLL-based builds, add self and bartowski to CODEOWNERS for tiled
* Put benchmarks behind a flag
* kernel fix for AVX2, iq quants
* Fix for asan, leaking memory in test-tiled-mulmat and avoid stack use after return
* guarding env flags with std::call_once
* Simplified repacking for VNNI to a single call per macrotile
* No threadlocals anymore, aligned wdata access
* Doing aligned reads since we ensure alignment with padding in wdata
* Eliminated per-thread gather of Q8_K rows in mul_mat_id, we now gather/repack in a single pass. Repack method now takes pointer array to support both dense/normal and mmid paths. Interface with ggml-cpu.c simplified as a result
* Unified/simplified dispatch and support checks. Put details on wdata needed inside the kernel.h body, simplified interactions with ggml-cpu.c.
* Cleanup includes and whitespace, update src1_repack to return false if we don't need a special repack, so the common case is handled by driver
* Better detection of win32 and additional whitespace fixes
* Gating fuzz tests behind a parameter and some extra prints to try and fix slow CI hosts
* Optimized AVX2 kernel
* Changed interleave format and added ability to interleave in-place after dequant
* Repacks now happen in-place, 16x64 microtiles are independent of each other
* Only repack rows in groups of 16 as they're needed. Save work in low n_rows cases and optimize L1 usage in other cases
* Use long panels for memory-bound regime (M <= 16), reintroduce IQP path for benchmarks
* Fix unused warnings and cleanup. Improved IQ dequantization speed.
* Removed separate process benchmarks
* Revert "Removed separate process benchmarks"
This reverts commit 0688cf43d5.
* AVX2 optimizations and guards for tests on windows
* Removed temp perf harness
* Remove perf-mulmat from build
* Removed IQP path, simplified tests to not use sub processes
* Cleaning up alignment of wdata
* Whitespace fixes and aligning L2 workspace to clean 512kb boundaries
* Update ggml/src/ggml-cpu/tiled/tiled-kernel.cpp
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
* Cleanup merge-duplicated declaration of test-backend-ops target
* Undo accidental line deletion in ggml.c
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
123 lines
6.3 KiB
Plaintext
123 lines
6.3 KiB
Plaintext
# collaborators can optionally add themselves here to indicate their availability for reviewing related PRs
|
|
# multiple collaborators per item can be specified
|
|
#
|
|
# ggml-org/ci : CISC, danbev, ggerganov, netrunnereve, ngxson, taronaeo
|
|
# ggml-org/ggml-cann : hipudding
|
|
# ggml-org/ggml-cuda : JohannesGaessler, am17an, IMbackK, ORippler
|
|
# ggml-org/ggml-hexagon : lhez, max-krasnyansky
|
|
# ggml-org/ggml-metal : ggerganov
|
|
# ggml-org/ggml-opencl : lhez, max-krasnyansky
|
|
# ggml-org/ggml-rpc : rgerganov
|
|
# ggml-org/ggml-sycl : arthw
|
|
# ggml-org/ggml-vulkan : 0cc4m, jeffbolznv
|
|
# ggml-org/ggml-webgpu : reeselevine, yomaytk
|
|
# ggml-org/ggml-zdnn : taronaeo
|
|
# ggml-org/llama-common : ggerganov, aldehir, angt, danbev, ngxson, pwilkin
|
|
# ggml-org/llama-mtmd : ngxson
|
|
# ggml-org/llama-server : ggerganov, ngxson, allozaur, angt, ServeurpersoCom
|
|
# ggml-org/llama-ui : allozaur
|
|
|
|
/.devops/*.Dockerfile @ngxson
|
|
/.github/actions/ @ggml-org/ci
|
|
/.github/workflows/ @ggml-org/ci
|
|
/ci/ @ggerganov
|
|
/cmake/ @ggerganov
|
|
/common/ @ggml-org/llama-common
|
|
/common/fit.* @JohannesGaessler
|
|
/common/jinja/ @CISC
|
|
/common/ngram-map.* @srogmann
|
|
/conversion/ @CISC
|
|
/convert_*.py @CISC
|
|
/docs/backend/snapdragon/ @ggml-org/ggml-hexagon
|
|
/examples/batched.swift/ @ggerganov
|
|
/examples/batched/ @ggerganov
|
|
/examples/convert-llama2c-to-ggml/ @ggerganov
|
|
/examples/debug/ @danbev @pwilkin
|
|
/examples/deprecation-warning/ @ggerganov
|
|
/examples/diffusion/ @am17an
|
|
/examples/embedding/ @ggerganov
|
|
/examples/eval-callback/ @ggerganov
|
|
/examples/export-docs/ @ggerganov
|
|
/examples/gen-docs/ @ggerganov
|
|
/examples/gguf/ @ggerganov
|
|
/examples/llama.android/ @ggerganov @hanyin-arm @naco-siren
|
|
/examples/llama.swiftui/ @ggerganov
|
|
/examples/llama.vim @ggerganov
|
|
/examples/lookahead/ @ggerganov
|
|
/examples/lookup/ @JohannesGaessler
|
|
/examples/model-conversion/ @danbev
|
|
/examples/parallel/ @ggerganov
|
|
/examples/passkey/ @ggerganov
|
|
/examples/retrieval/ @ggerganov
|
|
/examples/speculative-simple/ @ggerganov
|
|
/examples/speculative/ @ggerganov
|
|
/ggml/cmake/ @ggerganov
|
|
/ggml/include/ @ggerganov
|
|
/ggml/src/ggml-backend-meta.cpp @JohannesGaessler
|
|
/ggml/src/ggml-cann/ @ggml-org/ggml-cann
|
|
/ggml/src/ggml-common.h @ggerganov
|
|
/ggml/src/ggml-cpu/ @ggerganov
|
|
/ggml/src/ggml-cpu/tiled/ @jbooth @bartowski1182
|
|
/ggml/src/ggml-cpu/spacemit/ @alex-spacemit
|
|
/ggml/src/ggml-cuda/ @ggml-org/ggml-cuda
|
|
/ggml/src/ggml-cuda/vendors/hip.h @IMbackK
|
|
/ggml/src/ggml-hexagon/ @ggml-org/ggml-hexagon
|
|
/ggml/src/ggml-hip/ @IMbackK
|
|
/ggml/src/ggml-et/ @marty1885
|
|
/ggml/src/ggml-impl.h @ggerganov
|
|
/ggml/src/ggml-metal/ @ggml-org/ggml-metal
|
|
/ggml/src/ggml-opencl/ @ggml-org/ggml-opencl
|
|
/ggml/src/ggml-openvino/ @cavusmustafa @wine99
|
|
/ggml/src/ggml-opt.cpp @JohannesGaessler
|
|
/ggml/src/ggml-quants.* @ggerganov
|
|
/ggml/src/ggml-rpc/ @ggml-org/ggml-rpc
|
|
/ggml/src/ggml-sycl/ @ggml-org/ggml-sycl
|
|
/ggml/src/ggml-threading.* @ggerganov
|
|
/ggml/src/ggml-virtgpu/ @kpouget
|
|
/ggml/src/ggml-vulkan/ @ggml-org/ggml-vulkan
|
|
/ggml/src/ggml-webgpu/ @ggml-org/ggml-webgpu
|
|
/ggml/src/ggml-zdnn/ @ggml-org/ggml-zdnn @Andreas-Krebbel @AlekseiNikiforovIBM
|
|
/ggml/src/ggml-zendnn/ @avinashcpandey @Jiten1parmar @z-vishal
|
|
/ggml/src/ggml.c @ggerganov
|
|
/ggml/src/ggml.cpp @ggerganov
|
|
/ggml/src/gguf.cpp @JohannesGaessler @Green-Sky
|
|
/gguf-py/ @CISC
|
|
/media/ @ggerganov
|
|
/scripts/gen* @ggerganov
|
|
/scripts/get* @ggerganov
|
|
/scripts/sync* @ggerganov
|
|
/scripts/snapdragon/ @ggml-org/ggml-hexagon
|
|
/src/ @ggerganov
|
|
/src/llama-adapter.* @CISC
|
|
/src/llama-arch.* @CISC
|
|
/src/llama-chat.* @ngxson
|
|
/src/llama-graph.* @CISC
|
|
/src/llama-model.* @CISC
|
|
/src/llama-vocab.* @CISC
|
|
/src/models/ @CISC
|
|
/tests/ @ggerganov
|
|
/tests/test-chat.* @pwilkin
|
|
/tools/batched-bench/ @ggerganov
|
|
/tools/cli/ @ngxson
|
|
/tools/completion/ @ggerganov
|
|
/tools/mtmd/ @ggml-org/llama-mtmd
|
|
/tools/perplexity/ @ggerganov
|
|
/tools/parser/ @pwilkin
|
|
/tools/quantize/ @ggerganov
|
|
/tools/rpc/ @ggml-org/ggml-rpc
|
|
/tools/server/* @ggml-org/llama-server # no subdir
|
|
/tools/server/tests/ @ggml-org/llama-server
|
|
/tools/ui/ @ggml-org/llama-ui
|
|
/tools/tokenize/ @ggerganov
|
|
/tools/tts/ @ggerganov
|
|
/vendor/ @ggerganov
|
|
/AUTHORS @ggerganov
|
|
/CMakeLists.txt @ggerganov
|
|
/CONTRIBUTING.md @ggerganov
|
|
/LICENSE @ggerganov
|
|
/README.md @ggerganov
|
|
/SECURITY.md @ggerganov
|
|
/build-xcframework.sh @danbev
|
|
requirements*.txt @CISC
|
|
/skills @ngxson
|