cuda: support row-contiguous SUM_ROWS - #26308
Conversation
|
Hi @MohamedElashri, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
/bot review |
Automated code reviewReviewing the CUDA I read the PR description, the full CorrectnessThe change is correct. Verified points:
No blocking correctness or security issues. Will slow the review
Nits
Tests: This review was generated automatically by pi coding agent using |
There was a problem hiding this comment.
Thanks for the effort!
- Please add a test-case if possible (or confirm the tests already contained in the suite according to AI pass locally).
- IIRC, reduce_rows_f32 (i.e. the contiguous case) was used also for GGML_OP_MEAN. How much work would it be to extend this op to support row-contiguous tensors as well?
…ors using the same shared kernel, and add a test to MEAN permute/slice
Hi @ORippler, thanks 1- The existing I ran the following And the summary of results of the tests tried on RTX 3090 GPU (compute capability 8.6). I did some performance tests on explicit-stride synthetic cases for contiguous, permute-like, and slice-like row-contiguous layouts. Logical bandwidth below counts actual reduced row bytes plus output bytes. For sliced views this avoids inflating
If you got sometime, please review it and I would be happy to do extra work to finalize the best implementation. |
* upstream/master: (72 commits) HIP: Enable AllReduce for ROCm (ggml-org#27825) opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP (ggml-org#27637) ci: build MUSA for only 1 arch (ggml-org#28944) docs: Rule of thumb for AI review time [no ci] (ggml-org#28945) rpc : hash-cache only weights (ggml-org#28789) cuda: support row-contiguous SUM_ROWS (ggml-org#26308) models : move build_arch_graph() after graph() template specialization (ggml-org#28934) vulkan: support sparse Flash Attention (ggml-org#28105) OpenVINO: optimize stateful decode and GPU MoE inference (ggml-org#28638) opencl: add generic ssm_scan (ggml-org#28881) ci: bump kleidiai runners from 22.04 to 24.04 (ggml-org#28885) metal : add FA kernels for HSK=96, HSV=64 (MiniCPM3) (ggml-org#28599) ci: Bump CUDA Windows x64 builds to 13.4.1 (ggml-org#28930) ci : fix android release (ggml-org#28936) cuda : enable i16 and i32 for DUP (ggml-org#28897) cmake : use PROJECT_SOURCE_DIR instead of CMAKE_SOURCE_DIR (ggml-org#28771) webui: stop re-probing disabled /tools endpoint on every message (ggml-org#28646) ci : reuse build tag name when used instead of safe one (ggml-org#28911) CI: hip-quality-check: ignore spill added in bfdc321 (ggml-org#28909) HIP: fattn-mma: use fp32 accumulation on MFMA devices (ggml-org#28576) ...
* cuda: support row-contiguous SUM_ROWS * organize the code and add GGML_OP_MEAN to support row-contiguous tensors using the same shared kernel, and add a test to MEAN permute/slice * Keep original comments and add if/else branch
* cuda: support row-contiguous SUM_ROWS * organize the code and add GGML_OP_MEAN to support row-contiguous tensors using the same shared kernel, and add a test to MEAN permute/slice * Keep original comments and add if/else branch
* cuda: support row-contiguous SUM_ROWS * organize the code and add GGML_OP_MEAN to support row-contiguous tensors using the same shared kernel, and add a test to MEAN permute/slice * Keep original comments and add if/else branch
* cuda: support row-contiguous SUM_ROWS * organize the code and add GGML_OP_MEAN to support row-contiguous tensors using the same shared kernel, and add a test to MEAN permute/slice * Keep original comments and add if/else branch
Overview
This is straightforward, the PR extends CUDA
GGML_OP_SUM_ROWSfrom fully contiguous tensors to F32 row-contiguous tensors. What is does is that it adds a stride-aware kernel that derives the logicali1,i2, andi3coordinates for each row and calculates its address usingnb[1..3]. It also preserve the current behavio by using the existing reduction kernel as the fast path for fully contiguous tensors for the use cases it would be more suitable for.The validation/testing results is also complete (with focus on related stuff)
SUM_ROWS: 10/10SUM: 7/7MEAN: 7/7Requirements
Yes, AI was used. I used an agent to explore the codebase for codebase patterns and inspect possible style or merge-conflict problems. It also was used to see the test log files. But in both cases, I did re-check everything manually.