opencl: add generic ssm_scan - #28881
Merged
Merged
Conversation
lhez
marked this pull request as ready for review
September 14, 2026 16:52
max-krasnyansky
approved these changes
Sep 15, 2026
max-krasnyansky
left a comment
Member
There was a problem hiding this comment.
Nice!
We should have SSM_SCAN for ggml-hexagon up shortly as well.
dzannotti
added a commit
to halo-box/llama.cpp
that referenced
this pull request
Sep 15, 2026
* upstream/master: (72 commits) HIP: Enable AllReduce for ROCm (ggml-org#27825) opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (ggml-org#27637) ci: build MUSA for only 1 arch (ggml-org#28944) docs: Rule of thumb for AI review time [no ci] (ggml-org#28945) rpc : hash-cache only weights (ggml-org#28789) cuda: support row-contiguous SUM_ROWS (ggml-org#26308) models : move build_arch_graph() after graph() template specialization (ggml-org#28934) vulkan: support sparse Flash Attention (ggml-org#28105) OpenVINO: optimize stateful decode and GPU MoE inference (ggml-org#28638) opencl: add generic ssm_scan (ggml-org#28881) ci: bump kleidiai runners from 22.04 to 24.04 (ggml-org#28885) metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (ggml-org#28599) ci: Bump CUDA Windows x64 builds to 13.4.1 (ggml-org#28930) ci : fix android release (ggml-org#28936) cuda : enable i16 and i32 for DUP (ggml-org#28897) cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (ggml-org#28771) webui: stop re-probing disabled /tools endpoint on every message (ggml-org#28646) ci : reuse build tag name when used instead of safe one (ggml-org#28911) CI: hip-quality-check: ignore spill added in bfdc321 (ggml-org#28909) HIP: fattn-mma: use fp32 accumulation on MFMA devices (ggml-org#28576) ...
quimmedes
pushed a commit
to quimmedes/cafe-llama.cpp
that referenced
this pull request
Sep 16, 2026
* opencl: add generic ssm_scan * opencl: fix whitespace
zsogitbe
pushed a commit
to zsogitbe/llama.cpp
that referenced
this pull request
Sep 17, 2026
* opencl: add generic ssm_scan * opencl: fix whitespace
1 task done
wanghqc
added a commit
to qualcomm/llama.cpp
that referenced
this pull request
Sep 25, 2026
217 upstream commits since ad6c668, 8 of them in ggml-opencl. Two files conflicted, 16 hunks. - ssm_scan: upstream's generic kernel (ggml-org#28881) is added beside the specialised Mamba-2 kernels as the fallback for element-wise A and any other power-of-two d_state. The specialised d128/d256 kernels keep their row-folded variants and snapshot support; all of them are released when the device cannot give them a 64-lane subgroup, as upstream does. supports_op keeps the snapshot-width restriction and widens d_state to upstream's rule. - FA bin kernel (ggml-org#29046), q6_K ILA GEMM (ggml-org#28678, ggml-org#29057), q4_0/q4_K dp4a ILA GEMMs (ggml-org#29055, ggml-org#29056): taken as upstream wrote them. The capability check reads has_integer_dot_product, this branch's name for the field. - q4_0 MoE dp4a/ILA arbitration: kept this branch's version, which upstream adopted with the same routing threshold.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Current ssm_scan uses subgroups and assumes subgroup size is 64. It also requires d_state in {128, 256} and scalar A per head. This PR adds a more generic ssm_scan that loosens these conditions.
Additional information
Requirements