Skip to content

opencl: add bin kernel kernel_gemm_noshuffle_q4_0_q8_1_dp4a_ila_a8_bin - #29055

Merged
lhez merged 1 commit into
ggml-org:masterfrom
qualcomm:q4_0-a8-dp4a-bin-kernel
Sep 22, 2026
Merged

lhez merged 1 commit into
ggml-org:masterfrom
qualcomm:q4_0-a8-dp4a-bin-kernel

Conversation

@shaofeiqi

Copy link
Copy Markdown
Contributor

Overview

Add an optimized DP4A bin kernel for Q4_0 non-MoE GEMM.

Requirements

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend labels Sep 17, 2026
@lhez
lhez marked this pull request as ready for review September 22, 2026 02:20
@lhez
lhez requested a review from a team as a code owner September 22, 2026 02:20
@lhez
lhez merged commit ec5a12b into ggml-org:master Sep 22, 2026
24 of 28 checks passed
@BrewTestBot BrewTestBot mentioned this pull request Sep 23, 2026
1 task done
wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Sep 25, 2026
217 upstream commits since ad6c668, 8 of them in ggml-opencl. Two files
conflicted, 16 hunks.

- ssm_scan: upstream's generic kernel (ggml-org#28881) is added beside the
  specialised Mamba-2 kernels as the fallback for element-wise A and any
  other power-of-two d_state. The specialised d128/d256 kernels keep their
  row-folded variants and snapshot support; all of them are released when
  the device cannot give them a 64-lane subgroup, as upstream does.
  supports_op keeps the snapshot-width restriction and widens d_state to
  upstream's rule.
- FA bin kernel (ggml-org#29046), q6_K ILA GEMM (ggml-org#28678, ggml-org#29057), q4_0/q4_K dp4a
  ILA GEMMs (ggml-org#29055, ggml-org#29056): taken as upstream wrote them. The capability
  check reads has_integer_dot_product, this branch's name for the field.
- q4_0 MoE dp4a/ILA arbitration: kept this branch's version, which
  upstream adopted with the same routing threshold.
LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants