5127 Commits
Author SHA1 Message Date
Daniel Bevenius a2b36eb677 whisper : bump version to 1.9.4 (#4050) b5127 2026-09-10 12:17:05 +02:00
Daniel Bevenius 6fb4cd675c ci : add Windows On ARM support to release job (#4048)
* ci : add Windows On ARM support to release job (wip)

This commit adds Windows On Arm (WoA) support to whisper.cpp and the
release process.

The underlying support was already in place as this had been synced
with ggml, but the missing part was producesing release artifacts which
is what this commit does.

* ci : add note about CUDA WoA is a preview edition [no ci]
2026-09-10 09:51:35 +02:00
Javier De Jesus c44b60b805 whisper : call encoder_begin_callback before language auto-detect (#3936)
The language auto-detect path in whisper_full_with_state ran the encoder
without firing encoder_begin_callback, unlike the main transcription loop.
Callers that gate, observe, or abort whisper's encode through the callback
had no control over the auto-detection pass. Guard the auto-detect encode
with the callback the same way the main loop does.
2026-09-08 13:35:34 +02:00
davidcabcabanddavid 6d0ed91499 whisper : re-seed decoder 0 between calls (#4025)
Decoder 0 is seeded once in whisper_init_state() and the per-call re-seeding
loop in whisper_full_with_state() starts at j = 1, so decoder 0 is skipped.
Its mt19937 therefore carries over from one call to the next for the whole
lifetime of the state.

It is only consulted in the temperature > 0 branch (whisper_sample_token),
which is reached through temperature fallback -- so the effect only shows on
audio that falls back, and looks like nondeterminism rather than a bug: the
output becomes a function of how many calls the state has already served, and
the same audio decoded twice can yield different text.

Reproduced on a whisper-server instance: three noisy 30 s clips decoded four
times each in the same process gave 3/4 distinct transcripts per clip before
this change and identical transcripts after it. Fresh processes decoding each
clip once hide the issue entirely, which is why it is easy to miss.

Re-seeding decoder 0 with mt19937(0) matches what the loop does for every
other decoder (mt19937(j)) and makes decoding reproducible across calls.

Co-authored-by: david <david@cabrini.ch>
2026-09-08 11:58:05 +02:00
Daniel Bevenius cec4dbe5d5 ci : update close-issue job to not close issues (#4045)
This commit updates the close-issue CI job to not automatically close
issues. It also adds a stale lable message that explains why the issue
was marked as stale.

The motivation for this is that we enabled this job recently and there
have been a few comments from users/reporters that they did not get any
notification or motivation for closing. Hopefully marking the issues as
stale will give a notification the the reporters and we can manually
look through stale issues.

Refs: https://github.com/ggml-org/whisper.cpp/issues/586#issuecomment-5579674875
2026-09-08 11:57:20 +02:00
Abir Deol 79f2d92112 tests : load backends before init when built with GGML_BACKEND_DL (#4031)
With GGML_BACKEND_DL=ON no backend is registered until something calls
ggml_backend_load_all(). Every example does; the tests never did, so
whisper_init_* and parakeet_init_* ran with devices = 0 and aborted on
GGML_ASSERT(device) in ggml_backend_dev_backend_reg.

Three tests aborted on such a build - test-whisper-zero-samples,
test-vad and test-parakeet - taking ctest -L gh to 2 of 4.
test-vad-full has the same fault and reaches it only once a real
base.en model is present.

No workflow caught this because none pairs the two: build-clang,
build-gcc and build-sanitize run ctest -L gh but never set
GGML_BACKEND_DL, while release.yml sets it and runs no tests.

ggml_backend_load_all() is already reachable through whisper.h and
parakeet.h via ggml-cpu.h, so no new include is needed.

Resolves: https://github.com/ggml-org/whisper.cpp/issues/4030
2026-09-08 07:14:42 +02:00
Álvaro Justen 61e6ccade0 server : return language in detect response (#4035)
Resolves: https://github.com/ggml-org/whisper.cpp/issues/3603
2026-09-08 07:12:52 +02:00
Georgi Gerganov 52a939a2a7 sync : ggml 2026-09-04 13:40:00 +03:00
Georgi Gerganov a937f4e8ef ggml : bump version to 0.23.0 (ggml/1618) 2026-09-04 13:39:43 +03:00
Niklas Wenzel 11d4eec830 metal : add remaining fa-vec tunings for M3 Max (llama/28373) 2026-09-04 13:39:43 +03:00
Daniel Bevenius 140e57a4e2 ggml : replace compile definitions with version.h.in (llama/28364)
This commit adds a cmake version configuration file to replace the
current compile definition solution for the version.

The motivation for this change is that I made a mistake and did not take
into consideration that the compile definition means that this will
become a compiler flag for all sources in the target. This means that
when a version update happens that will recompile all sources in the
target even if they have not changed.

Refs: https://github.com/ggml-org/llama.cpp/pull/28278
2026-09-04 13:39:43 +03:00
Georgi Gerganov e2389eb99c ggml : rename and make private ggml_op_alloc_size_may_expand() (ggml/0)
cont https://github.com/ggml-org/llama.cpp/pull/27960
2026-09-04 13:39:43 +03:00
Adrien Gallouët 1b37beadd7 ggml : don't crash when backend search path can't be read (llama/28271)
Use std::error_code overloads of fs::current_path() and
fs::directory_iterator in ggml_backend_load_best() so an
inaccessible search path (WebDAV mount, removed CWD) is
skipped instead of terminating the process with an uncaught
filesystem_error.

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-04 13:39:43 +03:00
Adrien Gallouët f32e6fa0c8 ggml : remove GGML_CUDA_PEER_MAX_BATCH_SIZE (llama/28177)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-04 13:39:43 +03:00
Aaron Teo e1bbe40520 ggml-cpu(s390x) : fix q5_1 uninitialized v_acc (llama/28332)
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
2026-09-04 13:39:42 +03:00
Frosty40 d1e0e6491f sycl: fuse rms_norm+mul+add and add+add residual chains (llama/27610)
Fuse RMS_NORM+MUL+ADD and ADD+ADD under GGML_SYCL_ENABLE_FUSION.

ADD+ADD uses the same binbcast indexing and type matrix as standalone
add() (f32, f16, f16/f32, i32, i16, bf16, including broadcast and
non-contiguous). Unsupported combinations fall back to two add() launches.
2026-09-04 13:39:42 +03:00
Ozymandias_EBON 36f170e5fe SYCL: Refactor GGML_SYCL_ENABLE_MKL_FA to global var (llama/26863) 2026-09-04 13:39:42 +03:00
Hongqiang Wang d784add75f opencl: quant lm_head / decode GEMV and medium-batch GEMM optimizations (speculative decoding/MTP) (llama/26477)
* opencl: quant lm_head / decode GEMV and medium-batch GEMM optimizations

* opencl: guard q4_K/q6_K tiled_ns convert-kernel registration for non-Adreno build

* opencl: gate q4_K MUL_MAT+GLU fusion dispatch to Adreno

* opencl: require the noshuffle weight layout in the q4_K GLU fusion gate

* opencl: do not take the vectorized f16 mrow GEMV path on an unaligned row stride

* opencl: pass the new get_scale_min_k4 stride argument at the row-major call sites

* opencl: enable the q4_K split-K decode GEMV only where it is measured to win

* opencl: record the X1-85 split-K datapoint (neutral, exclusion confirmed)

* opencl: restrict the tiled lm_head/embed GEMV default to X2E/A8X

* opencl: fix q4_K variant kernels to read the transposed scales layout

* opencl: keep the flat-GEMV large-m escape opt-in

* opencl: guard the o4 GEMV store against the rounded-up dispatch tail

* opencl: restore the tiled q4_K/q6_K layout on tensor read-back

* opencl: split-K for the q8_0 decode GEMV at small M

* opencl: keep the q6_K noshuffle correctness escape ahead of the opt-in gate
2026-09-04 13:39:42 +03:00
kbenkhaled 0a4a95c86e tune MMVQ to MMQ crossover for SM87 (llama/28285) 2026-09-04 13:39:42 +03:00
Georgi Gerganov 4dd48dde35 metal : add sparse FA (llama/28098)
* metal : support n_kv_max sparse mask hint in flash attention vec kernel

- add kernel_flash_attn_ext_vec_idx: compacts finite mask entries into
  a per-row index list (Hillis-Steele scan, one threadgroup per row)
- extend vec FA kernel with optional sparse index gathering (FC slot 5)
- add host-side gate: sparse path when n_kv_max > 0, mask present,
  supported head sizes / KV types, n_kv_max <= 4096
- new buffer region extra_idx for the index list
- pipeline getter extended with has_sparse param
- add test cases: head sizes, quant types, nb>1, nr23 variants,
  sinks, ALiBi, softcap, permute, v_view_of_k, no-mask fallback

Note: multi-row (nb*nr23[1] > 1) cases still failing - rid mapping
in the store phase needs revisiting for the sparse path.

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* metal : fix sparse flash attention row addressing

- kernel_flash_attn_ext_vec_idx: mask param is half* but nb31 is a byte
  stride, so the per-row mask offset was scaled by 2x; cast to char*
  before applying the byte strides
- kernel_flash_attn_ext_vec: sparse pidx param is char* so the per-row
  element offset was under-scaled by sizeof(int); scale it by sizeof(int)
  to get the correct byte offset
- fixes the multi-row (nb*nr23[1] > 1) sparse flash attention failures

Assisted-by: pi:llama.cpp/DeepSeek-v4-0731

* cont : use sparse vec FA for prefill

* metal : single-pass flash attention sparse index compaction

The idx kernel previously read the mask row twice: once to count the finite
entries (for the prefix scan) and again to recover their positions. Since the
kernel is memory-bound, this doubled the mask traffic.

Keep the finite positions in a per-thread register array during the count
pass and write them out directly, avoiding the second mask read. A dense
mask with more than NLOCAL finite entries in a slice falls back to re-reading
the mask to write the remaining positions.

Assisted-by: pi:llama.cpp/DeepSeek-v4-0731

* tests : add perf cases for sparse flash attention prefill

Measure the sparse vec FA kernel across KV sizes, n_kv_max hints and batch
sizes. Run with:

    ./build/bin/test-backend-ops -b MTL0 -o FLASH_ATTN_EXT -p "n_kv_max=[1-9]" perf

Assisted-by: pi:llama.cpp/DeepSeek-v4-0731

* qwen4 : enable sparse attention

* cont : adjust nsg

* cont : sync test-backend-ops

* cont : disable Qwen4 for now

* cont : clean-up + tests
2026-09-04 13:39:42 +03:00
Georgi Gerganov d55d345e69 metal : fix glu dispatch with ne00 = 1 (llama/28306)
* metal : fix glu dispatch with ne00 = 1

* tests : disable ill-defined tests
2026-09-04 13:39:42 +03:00
25350b579e CUDA: Allow concurrent streams per split for multi-GPU (llama/28198)
* CUDA: Allow CUDA optimization per split for multi-GPU.

Previous guard caused multi-GPU to skip the graph optimization.  The
graph is already split per device and the optimization doesnt run
over the whole model but once per split, and thus should be allowed.
However, the CUDA event ggml_cuda_concurrent_event belongs to
whichever GPU was "current" when created. If the pass ran while
GPU 0 was current, it would stick and during event creation for the
second GPU it would land on GPU 0.

The fix: set the device explicitly ggml_cuda_set_device(cuda_ctx->device);
Default behaviour remains unchanged, only active for GGML_CUDA_GRAPH_OPT=1.
Explicit device setting pattern re-used from ggml_backend_cuda_graph_compute.

* Update ggml/src/ggml-cuda/ggml-cuda.cu

Co-authored-by: Aman Gupta <amangupta052@gmail.com>

---------

Co-authored-by: tannerbruhn <tannerbruhn@users.noreply.github.com>
Co-authored-by: Aman Gupta <amangupta052@gmail.com>
2026-09-04 13:39:41 +03:00
Nathan Wilson 47d348a278 vulkan: fix FA dequant path engagement (llama/28190)
Skip the nb[3] check when ne[3] == 1, the shader never reads it for a
single stream. Cache views carry the full-buffer stride there, so the old
check reduced to n_kv == kv_size and the path only engaged with the
cache full.
2026-09-04 13:39:41 +03:00
Neo Zhang f24a38605b sycl : enhance the api to support peer-to-peer copy (llama/27550) 2026-09-04 13:39:41 +03:00
EurekaticandRaulAbejonDelgado a704770e3b sycl: reduce redundant work in Q4_K multi-column MMVQ (llama/27062)
* sycl: Q4_K Weight unpack optimization and reuse between destination Columns

* sycl: Q4_K small N (N=2..4) + two output rows by subgroup reuse of activation between two rows.

* sycl: gate Q4_K two-row reuse for small N=2

* sycl: Fix on magic number now uses Q4_K_MMVQ_ROW_PAIR_MIN_NROWS=6272 for it, added tests for coverage around Q4_K_MMVQ_ROW_PAIR_MIN_NROWS with perf support to test Q4_K MUL_MAT, applied the same  reuse pattern to the activation as the weights.

Assisted-by: GPT-5.6 Sol

---------

Co-authored-by: RaulAbejonDelgado <raul.abejon.delgado@gmail.com>
2026-09-04 13:39:41 +03:00
Xuan-Son Nguyen e5605697ce finetune: fix no KV cache (llama/27199)
* training: fix no KV cache

* apply @ ggerganov
 suggestion
2026-09-04 13:39:41 +03:00
cqderek 37f0f443d7 ggml-hexagon: add F16 support for unary ops (llama/28228)
Extend the HTP backend's F16 unary op coverage to include ABS on top
of the existing NORM/RMS_NORM/L2_NORM/SCALE/CLAMP/SQR/SQRT set.

- Add hvx_abs_f16_{aa,au,ua,uu} + dispatcher in hvx-arith.h, mirroring
  the sqr_f16 kernel structure and using the existing hvx_vec_abs_f16()
  sign-bit-clear helper
- Add abs_f16() row-wise dispatch and DEFINE_UNARY_TASK_F16(unary_abs, ...)
  in unary-ops.c, wired into execute_op_unary()'s op_type/task_func
  switches
- Register HTP_OP_UNARY_ABS in htp_op_is_unary() (unary-ops.h) so that
  ggml_hexagon_precompute_unary_params() fills kernel_params (n_threads,
  VTCM layout) for ABS nodes -- required for the F16 path to function
- Narrow the F16 GGML_OP_UNARY gate in ggml_hexagon_supported_unary()
  (ggml-hexagon.cpp) to allow GGML_UNARY_OP_ABS specifically, instead of
  rejecting all GGML_OP_UNARY ops for F16
- Merge the separate execute_op_unary_f32()/execute_op_unary_f16()
  functions into a single execute_op_unary(), branching on an is_f16
  flag for the parts that actually differ by type (elem_size, the
  early F16 op-support check, and which task_func table to use) while
  keeping the F32-only tiled/RMS_NORM_MUL paths intact -- per review
  feedback to avoid duplicating the shared VTCM/DMA plumbing

Verified on-device (QRD8850, Hexagon v81) via test-backend-ops -o ABS:
8/8 passing (F16 + F32, HTP0, no CPU fallback). Regression-checked
SQR/CLAMP/SQRT (F16+F32) and NORM/RMS_NORM/L2_NORM/SCALE (F32; their F16
paths have no CPU reference kernel in test-backend-ops and cannot be
correctness-tested there independent of this change).
2026-09-04 13:39:41 +03:00
Isaac 1bdda1e366 metal : add fa-vec tunings for M3 (llama/28236) 2026-09-04 13:39:41 +03:00
3a1c7d6b6f metal : fix memory query under low-memory conditions (llama/27701)
* metal: Fix memory query under low-memory conditions

* Simply variable name

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Write it even shorter

Co-authored-by: Niklas Wenzel <dev@nikwen.de>

---------

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Co-authored-by: Niklas Wenzel <dev@nikwen.de>
2026-09-04 13:39:40 +03:00
Adrien Gallouët 4d343d7c05 ggml-cuda : remove unused vars (llama/28235)
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
2026-09-04 13:39:40 +03:00
Aman Gupta 519df618df CUDA + ggml: add sparse-fa for DSV4/GLM (llama/27970) 2026-09-04 13:39:40 +03:00
Aman Chadha(IVIXMMI)andAcmmi c2b400754b ggml: avoid KleidiAI buffer type init on dispatch (llama/27891)
Co-authored-by: Acmmi <acmmi@Acmmis-MacBook-Air.local>
2026-09-04 13:39:40 +03:00
Max Krasnyansky dc70853ec2 hexagon: MUL_MAT and MUL_MAT_ID fusion and fixes (llama/28202)
* hex-mm: fuse QKV and FFN matmuls that land on HMX

* hex-mm: remove hardcoded ne[1] < 32K restriction

* hex-get-rows: explicitly reject repacked Q8_0 just in case somebody decided to add an override

* hex-mm: correct overhead sizing to make sure we dont exceed vtcm budget for large dims

* hex-mm: fuse MUL_MAT_ID into MUL_MAT_ID_NX (2x,3x,...) where possible

* hex-fusion: update opbatch and opqueue sizing to acount for new fusion and reduce overhead for trace buffer alloc

* hex-bufs: sort buffers while finalizing opbatch, helps avoid va space fragmentation

* hex-bufs: add simple va defrag to make sure we dont abort just because the va space is fragmented

* hex-mm: replaced more scalar divs with fastdiv and minor cleanup

* hex-mm: tighten up supported fusion checks to exactly match supported kernels
2026-09-04 13:39:40 +03:00
Laurent ZuijdwijkandMarshall 1c7d35e14a vulkan: handle larger batch sizes (>4) efficiently for IQ3_S mat-vec (llama/27449)
* vulkan: handle larger batch sizes (>4) efficiently for IQ3_S mat-vec when NUM_COLS > 4. 5x perf at n=8

Assisted-by: Claude Opus 5

* adds 2 cases per quant type at `k=16*256` to the `all_types` mat-vec sweep

---------

Co-authored-by: Marshall <assistant@llama.cpp>
2026-09-04 13:39:40 +03:00
Mads Marquart a9e58612c5 vulkan : only request VK_KHR_shader_bfloat16 extension if supported (llama/28155) 2026-09-04 13:39:40 +03:00
Alan Tseng d57ae98246 ggml-cpu : conditionally add SpacemiT IME kernel sources (llama/27961)
When building with gcc < 15, ggml/CMakeLists.txt unconditionally adds
ime2_kernels.cpp, which fails to compile. FindSMTIME.cmake only defines
RISCV64_SPACEMIT_IME2 when the IME2 instructions are detected, and gcc 14
only has IME1, so ime2_kernels.cpp hits its #error.

This PR fixes it by using IN_LIST to add each kernel source according to
the spec that was actually detected.
2026-09-04 13:39:39 +03:00
Hongqiang Wang 35133c94c5 opencl: fix out‐of‐bound reads in the Adreno image kernels (#27632)
* opencl: clamp the q4_K decode GEMV's fetch row on a padded x-grid

* opencl: enforce the tiling contract of the image KQ/KQV GEMMs

* opencl: decide the image KQ/KQV split at the dispatch, not from strides
2026-09-04 13:39:39 +03:00
Trivikram Reddy c94921f8c6 hexagon: add missing FARF logs for cpy/get_rows/set_rows/gdn ops (llama/28217)
* hexagon: fix bug ne[2] printed in proc_op_req prep-src log

* hexagon: add shape/VTCM farf logs to cpy, get/set rows, gdn
2026-09-04 13:39:39 +03:00
Jhen-Jie Hong fcc2feee2f metal : add metallib build support for xcframework (llama/28163) 2026-09-04 13:39:39 +03:00
anujj 2c486783ab cuda: fuse MoE weighted expert reduction (llama/25952)
* cuda : fuse MoE weighted reduction (mul + view + add)

The MoE combine tail currently writes weighted expert outputs to
global memory before reducing them. That intermediate global-memory
traffic is the main cost. The production baseline generally runs two
physical fused kernels; this path runs one.

This change matches the full expert-weighting plus ordered-reduction
subgraph and replaces it with one weighted-reduction kernel.

Supported graphs:
- unscaled: experts * router_weights
- scaled:   (experts * expert_scale) * router_weights

k = 2..15 is handled by one runtime-k kernel.

Matching is structural: op sequence, shapes, strides, expert views,
and the left-to-right ADD chain. The fused kernel keeps that same
reduction order. Results are not claimed bit-identical; CUDA FP32
contraction can change rounding slightly.

Allocator integration uses add_alloc_dep from the graph-optimizer
API so experts, router weights, and optional expert scales stay live
until the fused destination is written. Memory ranges are rechecked
before the fused kernel runs.

Unrecognized or unsafe graphs are left alone and keep the existing
per-op path. Set GGML_CUDA_MOE_WEIGHTED_REDUCTION=0 to disable the
fusion.

test-backend-ops covers scaled/unscaled, aligned/unaligned, and
representative values across k=2..15, plus a k=16 case that must
stay on the per-op path.

* Pruned the test matrix from 15 to 6

* Addressed the aman and olivers review comments
2026-09-04 13:39:39 +03:00
Titaniumtown f162a19475 Revert "sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…" (#28184)
This reverts commit 1f3d318734c61cf6f3b209726cdd3f9c300a782e.
2026-09-04 13:39:39 +03:00
Jingxin (Philip) Li 408faaafce sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 1280 (llama/28016) 2026-09-04 13:39:38 +03:00
Lukasz Stolcman 5f07f856c5 metal : add fa-vec tuning for M2 Pro (llama/28122)
* metal: add fa-vec tuning for M2 Pro

* metal : update fa-vec tuning for M2 Pro with new dtypes
2026-09-04 13:39:38 +03:00
Jhen-Jie Hong 8cca1a3616 metal : add fa-vec tunings for A18 Pro (MacBook Neo) (llama/28152) 2026-09-04 13:39:38 +03:00
Niklas WenzelandYiChen Lv a245a8f48f metal : fix more leaks due to missing autoreleasepools (llama/27883)
* metal : fix more leaks due to missing autoreleasepools

* metal : rename variable

* metal : fix another missing pool warning

Co-authored-by: YiChen Lv <63285796+forforever73@users.noreply.github.com>

---------

Co-authored-by: YiChen Lv <63285796+forforever73@users.noreply.github.com>
2026-09-04 13:39:38 +03:00
Georgi Gerganov 870db2afa1 metal : add fa-vec tuning for M2 Max (llama/28015)
Rows for M2 Max (30 GPU cores) collected with 'ggml-metal-tuning fa-vec
--dtype f16,q8_0', pasted into fa_vec_tuned_table.

ref: https://github.com/ggml-org/llama.cpp/discussions/27668#discussioncomment-18205786

Assisted-by: pi:llama.cpp/Qwen3.8-27B
2026-09-04 13:39:38 +03:00
Neo Zhang 4f3a2a4b79 sycl : support limit max alloc memory within 2GB for host-pinned memory (llama/27559) 2026-09-04 13:39:38 +03:00
James Francis 5032008bc8 metal: enable Metal 4.0 tensor API on M5+/A19+ (llama/27461)
* metal : request Metal 4.0 language version for the tensor API

* metal : load the tensor API kernels from a separate metallib

* tests : add external-metallib tensor API regression test

* metal : fix metallib build order for the tensor API kernels
2026-09-04 13:39:38 +03:00
Buğra Özgürsoy 8e54c659b5 metal : add fa-vec tunings for M1 Ultra (llama/28088)
* metal : add fa-vec tunings for M1 Ultra

* metal : move M1 Ultra tunings after M1 Max section

* metal : remove duplicate blank line
2026-09-04 13:39:37 +03:00
ynankani f22bb2ea4d CUDA: XOR swizzle flash attn K,V smem fp16 tiles (llama/25635)
* CUDA: XOR swizzle flash attn  K,V smem fp16 tiles

Signed-off-by: ynankani <ynankani@nvidia.com>

* Fix use 64bit generic pointer instead of 32bit shared pointer

Signed-off-by: ynankani <ynankani@nvidia.com>

* fix shared memory race in FA on DGX Spark

* Handle corener case

Signed-off-by: ynankani <ynankani@nvidia.com>

* Add swizzle test cases and gate sync for swizzled path only

Signed-off-by: ynankani <ynankani@nvidia.com>

* gate CUDA PTX

Signed-off-by: ynankani <ynankani@nvidia.com>

* offset calculation specific for swizzle branch

Signed-off-by: ynankani <ynankani@nvidia.com>

* Reafctor code

Signed-off-by: ynankani <ynankani@nvidia.com>

* Refactor FA swizzle ldmatrix if/else into helpers (K row/col, V offset)

Signed-off-by: ynankani <ynankani@nvidia.com>

* rebase and update test case args

Signed-off-by: ynankani <ynankani@nvidia.com>

* Allow swizzle for non-pow2 shapes, for which nbatch_2%32==0

Signed-off-by: ynankani <ynankani@nvidia.com>

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-09-04 13:39:37 +03:00