Skip to content

opencl: add generic ssm_scan - #28881

Merged
ggerganov merged 2 commits into
ggml-org:masterfrom
qualcomm:lh/fix-ssm-scan
Sep 15, 2026
Merged

ggerganov merged 2 commits into
ggml-org:masterfrom
qualcomm:lh/fix-ssm-scan

Conversation

@lhez

@lhez lhez commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

Overview

Current ssm_scan uses subgroups and assumes subgroup size is 64. It also requires d_state in {128, 256} and scalar A per head. This PR adds a more generic ssm_scan that loosens these conditions.

Additional information

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes, asked codex to create the generic implementation and refactored several times.

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend labels Sep 14, 2026
@lhez
lhez marked this pull request as ready for review September 14, 2026 16:52
@lhez
lhez requested a review from a team as a code owner September 14, 2026 16:52

@max-krasnyansky max-krasnyansky left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice!
We should have SSM_SCAN for ggml-hexagon up shortly as well.

@max-krasnyansky max-krasnyansky added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 15, 2026
@ggerganov
ggerganov merged commit 6ec1a7e into ggml-org:master Sep 15, 2026
28 of 29 checks passed
dzannotti added a commit to halo-box/llama.cpp that referenced this pull request Sep 15, 2026
* upstream/master: (72 commits)
  HIP: Enable AllReduce for ROCm (ggml-org#27825)
  opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (ggml-org#27637)
  ci: build MUSA for only 1 arch (ggml-org#28944)
  docs: Rule of thumb for AI review time [no ci] (ggml-org#28945)
  rpc : hash-cache only weights (ggml-org#28789)
  cuda: support row-contiguous SUM_ROWS (ggml-org#26308)
  models : move build_arch_graph() after graph() template specialization (ggml-org#28934)
  vulkan: support sparse Flash Attention (ggml-org#28105)
  OpenVINO: optimize stateful decode and GPU MoE inference (ggml-org#28638)
  opencl: add generic ssm_scan (ggml-org#28881)
  ci: bump kleidiai runners from 22.04 to 24.04 (ggml-org#28885)
  metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (ggml-org#28599)
  ci: Bump CUDA Windows x64 builds to 13.4.1 (ggml-org#28930)
  ci : fix android release (ggml-org#28936)
  cuda : enable i16 and i32 for DUP (ggml-org#28897)
  cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (ggml-org#28771)
  webui: stop re-probing disabled /tools endpoint on every message (ggml-org#28646)
  ci : reuse build tag name when used instead of safe one (ggml-org#28911)
  CI: hip-quality-check: ignore spill added in bfdc321 (ggml-org#28909)
  HIP: fattn-mma: use fp32 accumulation on MFMA devices (ggml-org#28576)
  ...
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
* opencl: add generic ssm_scan

* opencl: fix whitespace
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* opencl: add generic ssm_scan

* opencl: fix whitespace
@BrewTestBot BrewTestBot mentioned this pull request Sep 23, 2026
1 task done
wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Sep 25, 2026
217 upstream commits since ad6c668, 8 of them in ggml-opencl. Two files
conflicted, 16 hunks.

- ssm_scan: upstream's generic kernel (ggml-org#28881) is added beside the
  specialised Mamba-2 kernels as the fallback for element-wise A and any
  other power-of-two d_state. The specialised d128/d256 kernels keep their
  row-folded variants and snapshot support; all of them are released when
  the device cannot give them a 64-lane subgroup, as upstream does.
  supports_op keeps the snapshot-width restriction and widens d_state to
  upstream's rule.
- FA bin kernel (ggml-org#29046), q6_K ILA GEMM (ggml-org#28678, ggml-org#29057), q4_0/q4_K dp4a
  ILA GEMMs (ggml-org#29055, ggml-org#29056): taken as upstream wrote them. The capability
  check reads has_integer_dot_product, this branch's name for the field.
- q4_0 MoE dp4a/ILA arbitration: kept this branch's version, which
  upstream adopted with the same routing threshold.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. OpenCL Issues specific to the OpenCL backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants