Skip to content

opencl: add bin kernel kernel_gemm_noshuffle_q6_k_f32_32b_trans_ila_a8_bin - #28678

Merged
max-krasnyansky merged 2 commits into
ggml-org:masterfrom
qualcomm:q6_k-a8-bin-kernel
Sep 18, 2026
Merged

max-krasnyansky merged 2 commits into
ggml-org:masterfrom
qualcomm:q6_k-a8-bin-kernel

Conversation

@shaofeiqi

Copy link
Copy Markdown
Contributor

Overview

Add optimized Q6_K non-MoE GEMM as bin kernel, plus the matching GEMV kernel for Adreno.

Requirements

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend labels Sep 10, 2026
@shaofeiqi
shaofeiqi marked this pull request as ready for review September 16, 2026 22:22
@shaofeiqi
shaofeiqi requested a review from a team as a code owner September 16, 2026 22:22
@lhez lhez added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 18, 2026
@max-krasnyansky
max-krasnyansky merged commit ec92815 into ggml-org:master Sep 18, 2026
27 of 29 checks passed
@BrewTestBot BrewTestBot mentioned this pull request Sep 23, 2026
1 task done
wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Sep 25, 2026
217 upstream commits since ad6c668, 8 of them in ggml-opencl. Two files
conflicted, 16 hunks.

- ssm_scan: upstream's generic kernel (ggml-org#28881) is added beside the
  specialised Mamba-2 kernels as the fallback for element-wise A and any
  other power-of-two d_state. The specialised d128/d256 kernels keep their
  row-folded variants and snapshot support; all of them are released when
  the device cannot give them a 64-lane subgroup, as upstream does.
  supports_op keeps the snapshot-width restriction and widens d_state to
  upstream's rule.
- FA bin kernel (ggml-org#29046), q6_K ILA GEMM (ggml-org#28678, ggml-org#29057), q4_0/q4_K dp4a
  ILA GEMMs (ggml-org#29055, ggml-org#29056): taken as upstream wrote them. The capability
  check reads has_integer_dot_product, this branch's name for the field.
- q4_0 MoE dp4a/ILA arbitration: kept this branch's version, which
  upstream adopted with the same routing threshold.
LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
…a8_bin` (ggml-org#28678)

* opencl: add A8 Q6_K non-MoE binary kernel

* opencl: fix layout compatibility
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…a8_bin` (ggml-org#28678)

* opencl: add A8 Q6_K non-MoE binary kernel

* opencl: fix layout compatibility
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. OpenCL Issues specific to the OpenCL backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants